{"id":"3f2b3b6f-e79e-4863-8a29-236ba72f9a3f","arxiv_id":"2504.21051","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review cataloging medical MLLMs, their uses in report generation, diagnosis, and treatment, and the challenges of accuracy, hallucination, fairness, privacy, and deployment.","lead":"This paper surveys 330 recent publications on multimodal large language models (MLLMs) applied to medicine, organizing them into applications, data types, and open challenges. It offers a compact map of a fast-moving field for researchers and clinicians who want a quick orientation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Duplicate references and mis-citations (e.g., HuatuoGPT→[19], Med-PaLM as image-capable) undercut the '330 separate papers' comprehensiveness claim.","rationale":"The reader's CONDITIONAL verdict identifies the same general weakness: the survey's value depends on the representativeness and fidelity of its 330-paper synthesis, and the paper provides no systematic selection or verification methodology. My stress-test agrees with that framing but sharpens it: the manuscript contains direct internal evidence that the synthesis is not reliable, including duplicate references that inflate the paper count, at least one table row citing the wrong paper (HuatuoGPT→[19]), and a factual error about Med-PaLM's modality. These are concrete, checkable failures rather than merely an absent methodology. They reinforce the reader's condition that the survey needs a corrected, verifiable reference corpus and accurate per-paper descriptions before it can be considered a dependable map. I do not see a reason to move the verdict to ACCEPT or REJECT: the core narrative about medical MLLMs is broadly reasonable and the errors are localized and fixable, but the stated 'comprehensive review of 330 papers' cannot stand as-is. Hence the reader's CONDITIONAL verdict remains appropriate, and my concern does not change it.","tokens_in":45067,"tokens_out":4236,"duration_ms":40000,"concrete_test":"Build the unique-reference set from the bibliography (deduplicating by title/DOI), and compare the count against the claimed 330. Then independently fetch the full texts or abstracts of 20 randomly sampled entries from Tables 1 and 6 and check each row's stated base model, modality, and task description against the cited source. If the unique count falls below 330 or more than 3 of 20 sampled descriptions materially disagree with the cited paper, the '330-paper comprehensive review' claim must be revised to state the actual included corpus and the inclusion criteria used.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central value is its claim to be a comprehensive, reliable map of 330 recent medical-MLLM papers (Abstract, Section 3). That claim depends on the referenced papers being distinct, correctly described, and fairly synthesized. The manuscript's own internal evidence contradicts this. (1) The bibliography contains duplicate entries: [139] and [140] are identical, as are [145]/[146], [207]/[210], [126]/[127], and [4]/[221]. The true unique count is therefore less than 330, so the headline number is not trustworthy. (2) Table 6 attributes 'HuatuoGPT' to reference [19], which is actually the HuatuoGPT-Vision paper (arXiv:2406.19280), not the original HuatuoGPT; the same row lists BLOOMZ as the base model, which does not match the original HuatuoGPT's LLaMA base. (3) Table 6 describes Med-PaLM (a text-only LLM for medical QA) as having 'exceptional performance particularly in areas such as clinical note summarization, radiological image analysis, and drug interaction prediction,' but Med-PaLM does not process images. (4) The abstract promises three application directions (medical reporting, diagnosis, treatment) and 'six mainstream modes of data,' but Section 3 actually covers report generation, communication, and surgery assistance, and Section 4 presents only four data modalities (image, text, audio, omics). These are not mere stylistic slips: they indicate that the secondhand descriptions and the paper-counting that anchor the 'comprehensive review' claim have not been systematically verified. A survey is only as good as its sourcing, and these errors make the central claim unverified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of multimodal large language models (MLLMs) applied to medicine. It opens with background on LLMs and MLLMs, then discusses applications (claimed in the abstract to be medical reporting, diagnosis, and treatment), data modalities, model traits (professionalism, hallucination, fairness), and future challenges. The paper states that the review is based on 330 recent papers and that it covers six mainstream data modes with corresponding evaluation benchmarks. The body includes large tables of medical MLLMs, medical LLMs, training datasets, and descriptions of representative systems.","tokens_in":45403,"tokens_out":5440,"duration_ms":49285,"significance":"If the survey's content were accurate and its bibliographic claims verifiable, the manuscript would be a useful entry point for researchers entering medical MLLMs: the model tables, dataset lists, and the discussion of traits such as hallucination and fairness are organized in a readable way. The paper does not claim new algorithmic results, but the value of a survey lies in the fidelity of its summarization and the reliability of its coverage. At present, internal evidence undercuts both: the reference list contains multiple duplicated entries, the abstract advertises a structure that the body does not follow, and key system descriptions in Table 6 are mis-cited. These are not peripheral stylistic issues; they affect the central claim of being a comprehensive and reliable map of 330 papers.","major_comments":[{"comment":"The abstract's claim of a \"comprehensive review of 330 recent papers\" is not supported by the bibliography. The reference list contains multiple duplicated entries: [139] and [140] are the same multi-omics integration paper, [145] and [146] are the same medical-textbook reasoning paper, [207] and [210] are the same training-data extraction paper, [126] and [127] are the same MedDialog paper, and [4] and [221] are the same Vicuna reference. Additional duplicates include [102]/[109], [63]/[128]/[265], [22]/[273], [31]/[71], and [33]/[112]. The unique number of distinct works reviewed is therefore smaller than 330, and the headline number is not trustworthy as stated.","section":"References and Abstract"},{"comment":"The abstract promises three application directions: medical reporting, medical diagnosis, and medical treatment. Section 3, however, is organized around medical report generation (Section 3.1), professional and compassionate medical communication (Section 3.2), and clinical surgery assistance (Section 3.3). No section is devoted to medical diagnosis or medical treatment as such, so the advertised structure does not match the content actually presented.","section":"Section 3 versus Abstract"},{"comment":"The abstract states that the paper presents \"six mainstream modes of data along with their corresponding evaluation benchmarks.\" Section 4 details only four data categories: image (Section 4.1), text (Section 4.2), and two subsections under \"Data with Potential\" (audio in Section 4.3.1 and omics in Section 4.3.2). Moreover, no evaluation benchmarks are systematically paired with each modality in that section. The \"six modes\" claim must either be implemented in the text or removed from the abstract.","section":"Section 4 versus Abstract"},{"comment":"Table 6 contains load-bearing mis-citations. The HuatuoGPT row cites reference [19], which is the HuatuoGPT-Vision paper (arXiv:2406.19280), and lists BLOOMZ as the base model; the original HuatuoGPT is a text-only model built on LLaMA, not BLOOMZ. In the same table, the Med-PaLM row attributes to Med-PaLM \"exceptional performance particularly in areas such as clinical note summarization, radiological image analysis, and drug interaction prediction.\" Med-PaLM is a text-only medical question-answering model and does not process images, so this description is inaccurate.","section":"Table 6"},{"comment":"Table 5 is captioned \"Medical Multimodal Large Language Model Training Data,\" but its rows list text-only datasets such as MedDialog, Huatuo-26M, PubMedQA, and MIMIC-III. These are not multimodal training data. Table 4 is the actual multimodal dataset table, making the Table 5 heading misleading and in need of correction.","section":"Table 5"},{"comment":"The paper does not describe a search strategy, inclusion or exclusion criteria, or any protocol for selecting the claimed 330 papers. Without such methodology, the claim of comprehensiveness cannot be independently assessed, and the reader cannot tell whether the curated set is representative of the field or skewed toward particular topics, venues, or time periods.","section":"Sections 3 and 4 (methodology)"}],"minor_comments":[{"comment":"The word \"Moereover\" should be \"Moreover.\"","section":"Section 3.3"},{"comment":"The model names \"Med-Palm\" and \"Med-Palm M\" should be written consistently as \"Med-PaLM\" and \"Med-PaLM M,\" and there is an extra closing parenthesis after \"HuatuoGPT [19].\"","section":"Figure 1"},{"comment":"The model name \"Aqulia-Med\" should be \"Aquila-Med,\" and the rows are not in chronological order (for example, the 2023/11 MEDITRON row appears after 2024/01 and 2024/02 rows).","section":"Table 6"},{"comment":"Beyond the duplicates listed in the major comments, the reference list should be deduplicated globally; examples include [63], [128], and [265], which all point to the same radiology foundation-model paper.","section":"References"},{"comment":"The row for \"SAT-DS\" cites reference [257], which appears to be a paper about universal medical image segmentation rather than a question-answering dataset; this citation should be verified and corrected.","section":"Table 4"},{"comment":"Reference [54] is formatted as \"GPT-3(text-davinci-003)\" but the cited work is the InstructGPT paper \"Training language models to follow instructions with human feedback\"; the label should be made consistent with the cited work.","section":"Reference [54]"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is not acceptable in its current form because the central bibliometric claim (330 distinct papers) is contradicted by the reference list itself, and several key system descriptions in Table 6 are inaccurate. These issues are fixable in revision through a careful deduplication pass, systematic verification of every model-to-reference mapping, and alignment of the abstract with the actual structure of Sections 3 and 4. I would not reject outright because the qualitative organization has value, but the survey's reliability as a reference work depends on correcting these problems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You can trust the survey's organization more than its scholarship. It's a reasonable first map of medical MLLM applications, but the \"330 papers\" claim is inflated by duplicate references, and the abstract's \"six data modes\" is not what Section 4 delivers.\n\nWhat the paper does well: it collects a broad set of recent work and structures it in three application chapters (report generation, communication, surgery) plus separate discussions of data modalities and model-level challenges. The tables grouping models by base LLM and dataset types are convenient for orientation. A newcomer to medical MLLMs will finish with a decent sense of the landscape.\n\nThe soft spots are not cosmetic. The bibliography contains multiple duplicate entries: [139]/[140], [145]/[146], [207]/[210], [126]/[127], [4]/[221], and also [34]/[63]/[128] are the same paper. That means the headline \"330 papers\" is not a verified count. Table 6 misattributes HuatuoGPT to reference [19] — which is actually HuatuoGPT-Vision — and lists BLOOMZ as the base model, which the original HuatuoGPT did not use. The same table credits Med-PaLM with radiological image analysis; Med-PaLM is text-only. Table 5 is labeled \"Multimodal\" but lists text-only LLM datasets. And the abstract promises six mainstream data modes while Section 4 presents only image, text, audio, and omics. There is also no systematic search or inclusion methodology, so the selection of papers is unverifiable.\n\nThese flaws matter because a survey's value depends on the fidelity of its secondhand descriptions. If the descriptions and counts can't be trusted, the \"comprehensive\" claim weakens. The underlying structure is salvageable, and the writing is generally clear, but the reference list and internal consistency need a serious audit.\n\nWho this is for: someone looking for a quick orientation to medical MLLMs, with the caveat that they should check primary sources. I wouldn't cite it as a reliable map in my own work. If it goes to peer review, I'd send it to a referee but require the authors to fix duplicate citations, correct the Table 6 mis-citations, add a methodology section, and reconcile the abstract with the actual content. Without those changes, I would not accept it.","headline":"A useful but unreliable map of medical MLLMs: the broad organization is sound, but the paper count is inflated by duplicate references, the abstract overpromises six data modes, and several key citations are wrong.","tokens_in":45875,"tokens_out":3427,"would_cite":false,"duration_ms":35486,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to provide a comprehensive map of medical multimodal large language models built from 330 papers, organized around three clinical applications and six data modes.","keywords":["multimodal large language models","medical artificial intelligence","survey","medical report generation","medical dialogue","surgical assistance","medical benchmarks","clinical challenges"],"falsifier":"Run a documented literature search for medical vision-language models published through early 2025, draw a random sample of in-scope papers, and check whether they appear in the survey's tables; if a substantial share, say 20 percent or more, is missing, the 330-paper map is incomplete. Also spot-check descriptions by re-reading the cited originals and comparing specific claims such as state-of-the-art scores on named benchmarks.","tokens_in":44913,"feed_emoji":"🩺","tokens_out":6410,"duration_ms":62456,"temperature":0.7,"pith_summary":"This survey claims to provide a comprehensive map of multimodal large language models (MLLMs) in medicine, built from 330 recent papers. It organizes the field into three application directions: generating medical reports, conducting professional and compassionate medical communication, and assisting clinical surgery. For each application it identifies representative models, the data modalities they consume, and the benchmarks used to evaluate them. The survey also catalogs data types, including image, text, audio, omics, and hybrid forms, and argues that data scarcity is the main bottleneck. It concludes that medical MLLMs show promise but still need better evaluation benchmarks, handling of rapidly updated knowledge, edge deployment, and privacy safeguards before clinical use.","feed_headline":"330 medical multimodal AI papers mapped by clinical role","feed_subtitle":"A new taxonomy sorts the field into reporting, diagnosis, and treatment, with the data and benchmarks behind each.","key_machinery":"The load-bearing object is the standard MLLM architecture: a pretrained LLM at the core, a modality encoder for inputs such as images, video, or audio, an alignment module that projects those other modalities into the language model's feature space, and a generative model at the output. The survey uses this architectural recipe as its lens for classifying all 330 papers, and pairs it with a data taxonomy (image, text, audio, omics, instruction-following, and hybrid) and a list of evaluation benchmarks to explain how each model acquires and demonstrates medical competence. The taxonomy itself is the other central machinery: it is what turns a scattered literature into an ordered map.","core_discovery":"The paper's central claim is that the medical MLLM landscape can be read through a single taxonomy: three clinical application areas, six mainstream data modes, and corresponding evaluation benchmarks, all resting on one architectural recipe. That recipe places a pretrained large language model at the core, with modality-specific encoders at the input and an alignment module that fuses non-text features into the language model's feature space. The survey's tables list dozens of medical MLLMs and LLMs, their base models, and their training data, and it presents the main evaluation routes: text-similarity metrics, expert manual scoring, AI-as-judge scoring, and medical licensing examinations such as USMLE. On this evidence the paper argues that while MLLMs can score well on closed benchmarks, they remain short of what clinical practice requires, and it names professionalism, hallucination, fairness and bias, rapidly changing medical knowledge, deployment constraints, and privacy as the barriers to close.","pith_inferences":["Because the survey does not document a systematic search strategy, its map is best read as a literature snapshot; a protocol-driven search could test whether 330 papers captures the full population.","The paper's own examples imply that exam-style benchmarks such as USMLE overestimate clinical readiness; building evaluation suites from real clinician workflows would test this directly.","Since most surveyed models fine-tune general-purpose MLLMs, the field's progress likely tracks general multimodal releases; evaluations should be re-run when new foundation models appear.","Treating audio and omics as data with potential suggests a concrete testable program: add respiratory audio or genomic data to text-plus-imaging models and measure whether fusion improves diagnostic accuracy over single-modality baselines."],"forward_implications":["A newcomer can locate any medical MLLM paper under one of three headings, namely report generation, medical communication, or surgical assistance, and immediately see which data and benchmarks that line of work uses.","The six data-mode classification makes explicit that audio and omics are the least developed, so future data-collection efforts should concentrate there.","Because the architectural recipe uses a pretrained LLM at the core, medical MLLMs largely inherit their reasoning from general multimodal models, so gains in general pretraining should transfer to medicine.","The evaluation checklist provides a standard way to compare new models, and the survey's examples show why exam scores alone cannot certify clinical readiness.","The challenges named, hallucination, bias, privacy, deployment, and rapidly changing knowledge, define a concrete agenda for turning promising models into clinically usable systems."],"supporting_citations":[{"why":"Supplies the Transformer and self-attention architecture that the MLLMs in this survey all build on.","marker":"[1]"},{"why":"Serves as the model whose release drove medical multimodal research and as a baseline in medical evaluations.","marker":"[15]"},{"why":"Provides the medical LLM baseline used in the survey's evaluation discussion and in licensing-exam comparisons.","marker":"[17]"},{"why":"Shows how GPT-4V can denoise and reformat data, and contributes a medical vision-language model used as an example.","marker":"[19]"},{"why":"Supplies a medical multimodal few-shot learner and a key data point in the model tables.","marker":"[21]"},{"why":"Flagship medical vision-language assistant used throughout the survey for image-text tasks and benchmarks.","marker":"[22]"},{"why":"Represents the generalist biomedical MLLM direction and backs the claim that multimodal medical models can handle varied inputs.","marker":"[23]"},{"why":"Demonstrates a medical MLLM applied to skin imaging and interactive diagnosis, illustrating the medical communication application.","marker":"[24]"},{"why":"Provides the PMC-VQA dataset and model used to train and evaluate medical visual question answering.","marker":"[58]"}],"fun_headline_variants":["Medical AI survey: 330 papers, one taxonomy","MLLM medicine: reporting, diagnosis, treatment mapped","Medical MLLMs: from benchmarks to bedside gaps","Survey: medical AI needs more than good test scores","One taxonomy to map 330 medical AI papers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that its selected 330 papers are representative of the medical MLLM literature and that its secondhand descriptions of them are accurate, yet it does not document a search protocol or inclusion criteria.","fun_headline_variants_meta":{"raw":{"variants":["Medical AI survey: 330 papers, one taxonomy","MLLM medicine: reporting, diagnosis, treatment mapped","Medical MLLMs: from benchmarks to bedside gaps","Survey: medical AI needs more than good test scores","One taxonomy to map 330 medical AI papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3217,"prompt_tokens":914,"completion_tokens":2303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":2227}},"tokens_in":530,"tokens_out":2303,"duration_ms":17583,"temperature":1.0,"reasoning_tokens":2227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:29:16.652657+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a documented literature search for medical vision-language models published through early 2025, draw a random sample of in-scope papers, and check whether they appear in the survey's tables; if a substantial share, say 20 percent or more, is missing, the 330-paper map is incomplete. Also spot-check descriptions by re-reading the cited originals and comparing specific claims such as state-of-the-art scores on named benchmarks.","supporting_citations":[],"review_version":1}