{"id":"642da26c-75a5-4a4c-bea3-f232c66c1604","arxiv_id":"2505.06945","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 69 medical multimodal modeling studies, organized by five reported challenges and the solution families proposed for each.","lead":"This paper is a systematic review of 69 studies that model several types of medical data together, such as images, genomes, and electronic health records. It groups the difficulties researchers report, like missing data and small samples, and lists the methods used to address them.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cited solution families come partly from studies outside the claimed 69-study corpus (e.g., refs. [96] and [102] in the interpretability and optimal-fusion sections), so the challenge-solution map is not strictly derived from the reviewed set.","rationale":"The reader's weakest assumption was that the PubMed/Google Scholar search yields a representative corpus. That is a plausible external-validity limitation. The concern I identify is an internal-consistency problem: the manuscript cites at least two studies as if they were part of the 69-study corpus when they are not listed in Table 3, and one ([96]) is from 2024, after the search date. A systematic review's central value is the reproducibility and integrity of its corpus; if the corpus description is inaccurate, the challenge-to-solution counts and gap analysis are not trustworthy as stated. This is a stronger, more specific objection than search representativeness because it does not depend on an argument about what the search might have missed; it is verifiable from the text itself. It does not overturn the overall utility of the review, but it reinforces the CONDITIONAL disposition: the authors must reconcile the corpus list with the text and the supplementary extraction table. The reader's rationale already noted citation errors generally; my concern identifies the specific, load-bearing instances and shows how they affect the central claim.","tokens_in":25122,"tokens_out":8983,"duration_ms":79051,"concrete_test":"Extract the full list of included studies from Table 3 and Supplementary Table S1. Check whether references [96], [102], [94], and [81] are present. If [96] and [102] are absent, re-read the 'Interpretability' and 'Optimal fusion technique' subsections and determine whether their methods are presented as part of the reviewed corpus. Then either (a) add these studies to the inclusion list and recompute the challenge counts (Table 2, text totals), or (b) recategorize them as external context and remove/relabel the corresponding sentences. If the totals '14 publications' (interpretability) or '13 studies' (optimal fusion) change, or if the examples are not in the corpus, the claim that the 69 studies provide the reported solution families is not supported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that five recurring challenges and their solution families are identified from 69 systematically reviewed studies. That claim requires every solution attributed to the corpus to actually be one of the 69. The manuscript violates this in at least two places. In the 'Interpretability' section, Jiang et al. [96] (AUTOSurv, npj Precision Oncology 2024) is discussed together with the included Hao et al. [30] as addressing the black-box problem in survival analysis; [96] is not in Table 3's list of 69 studies and postdates the search cutoff (October 2023). In the 'Optimal fusion technique' section, after stating 'A total of 13 studies in our review identified finding the optimal fusion strategy as a challenge,' the text introduces Xu et al. [102] (MUFASA) as an example; [102] also does not appear in Table 3. If these references are not in the 69-study extraction, then the review's solution families draw on external work while being presented as findings of the reviewed corpus. This makes the prevalence counts and the gap analysis less reliable: the map is not solely a synthesis of the 69 studies claimed, and it breaks the internal consistency of the systematic review. The issue is directly checkable from the paper and Supplementary Table S1.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This systematic review claims to synthesize 69 studies on modeling multimodal medical data, using a PubMed search plus a manual Google Scholar sweep through late 2023. The authors identify five recurring technical challenges—missing modalities, small datasets, interpretability, imbalance in dimensionality, and selection of the optimal fusion strategy—and map reported solutions to each challenge, including transfer learning, generative models, attention mechanisms, knowledge distillation, and neural architecture search. The synthesis is organized around a challenge-to-solution table (Table 3) and supplemented by distributions of fusion strategies, medical domains, modality combinations, and performance metrics. The paper positions itself as the first challenge-oriented systematic map of methods for multimodal medical modeling.","tokens_in":25397,"tokens_out":4283,"duration_ms":43681,"significance":"If the synthesis is reliable, the paper has practical value as a challenge-oriented reference for practitioners and a gap analysis for researchers, complementing existing modality- or fusion-centric reviews. The authors provide a reproducible search query (Table A1), a dual-screening protocol, and a per-study extraction table, and they state limitations about search coverage and the deep-learning skew of the retrieved literature. The main quantitative claims, however, depend on the representativeness of the 69-study corpus and on the internal consistency of the PRISMA accounting; both need correction before the review can be used as a trustworthy map of the field.","major_comments":[{"comment":"The PRISMA flow diagram does not reconcile with the reported inclusion count. Figure 2 shows 84 records screened full-text, 82 reports assessed for eligibility after 2 reports were not retrieved, 5 studies removed (3 video recordings, 2 retracted papers), and 24 studies excluded, yielding 58 included. The arithmetic is 82 − 5 − 24 = 53 (or 84 − 2 − 5 − 24 = 53), not 58. Please clarify at which stage the 5 removed studies were excluded and correct the diagram or the counts, because the central claim \"69 studies\" depends on this accounting.","section":"Methods/Results, Figure 2"},{"comment":"Solution families attributed to the reviewed corpus are partly drawn from studies that are not among the 69 included studies. In the \"Interpretability\" section, Jiang et al. [96] (AUTOSurv) is discussed together with the included Hao et al. [30] as addressing the black-box problem in survival analysis, but [96] is absent from Table 3 and postdates the October 2023 search cutoff. In the \"Optimal fusion technique\" section, after the statement \"A total of 13 studies in our review identified finding the optimal fusion strategy as a challenge,\" the text introduces Xu et al. [102] (MUFASA) as an example, and [102] is also absent from Table 3. To preserve the internal consistency of the systematic review, these references must either be explicitly labeled as external context or be removed from the prevalence-based narrative; the challenge counts and the gap analysis should be recomputed from the included studies only.","section":"Challenges and solutions, Interpretability and Optimal fusion technique"},{"comment":"The search design strongly shapes the prevalence counts that are presented as descriptive of the field. The PubMed query requires the word \"challenges\" (line 4 of Table A1) and title-level occurrence of \"model fusion,\" \"data fusion,\" or \"multimodal\" (line 1), and the manual Google Scholar sweep is not documented with the same transparency. Studies that address missing modalities, small data, or fusion selection without framing them as \"challenges,\" or that use other vocabulary in the title, are systematically excluded. The reported percentages such as 15/69 and 17/69 are therefore prevalence within a query-defined corpus, not prevalence in the literature. Please either add a sensitivity analysis without the \"challenges\" restriction or explicitly reframe the quantitative claims as descriptive of the retrieved set rather than of the field.","section":"Methods, Search strategy and Table A1"}],"minor_comments":[{"comment":"The roadmap sentence says the remainder is organized with Section 6 for methodology, Section 6 for challenges, Section 6 for discussion, and Section 6 for conclusion; these placeholders should be replaced with the actual section numbers.","section":"Introduction, last paragraph"},{"comment":"Reference [81] appears to be a duplicate of reference [74] (both list Xu et al., \"Explainable Dynamic Multimodal Variational Autoencoder for the Prediction of Patients With Suspected Central Precocious Puberty\"); one entry should be removed and the in-text citations adjusted.","section":"References"},{"comment":"The four \"Studies excluded\" categories each show n=6, but two of the labels (\"Do not address Modeling challenge\" and \"Do not address any challenge\") are nearly indistinguishable; clarify the distinction so the exclusion flow is interpretable.","section":"Figure 2"},{"comment":"The modality-combination counts in Table 4 and Figure 4 should be cross-checked for consistency; for example, the triple combination \"Imaging + Genomic + Clinical\" is shown with count 3 in the figure while the table groups studies in a way that is easy to misread. Please align the two displays or add an explicit note on how multi-task rows are counted.","section":"Table 4 / Figure 4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this systematic review gives a reasonable challenge-to-solution map for multimodal medical data modeling, organized around five recurring technical challenges. The taxonomy is sensible and the paper is honest about being a synthesis, not a source of new methods. But the paper has an internal-consistency problem you should know about before citing it: at least two examples of \"studies in our review\" are not among the 69 included studies. In the interpretability section, Jiang et al. [96] (AUTOSurv, npj Precision Oncology 2024) is discussed alongside Hao et al. [30] as addressing the black-box problem, but [96] postdates the October 2023 search cutoff and is not in Table 3. In the optimal-fusion section, Xu et al. [102] (MUFASA) is presented as an example from the review, but [102] is also absent from Table 3. That breaks the claim that the solution families are strictly derived from the reviewed corpus. It is easily fixable—either drop those examples or add them with justification—but as written, the prevalence counts and gap analysis rest on a corpus that is not exactly the one claimed.\n\nWhat the paper does well: the five categories (missing modalities, small data, interpretability, dimensionality imbalance, optimal fusion) cover the landscape well, and the per-study table (Table 3) gives practitioners a quick entry point. The authors explicitly acknowledge search limitations, including that most included work is deep learning by result rather than design. That is good practice.\n\nWhere it is soft beyond the [96]/[102] issue: the PRISMA flow diagram does not account for the five removed studies (three video, two retracted) in the 84-to-58 arithmetic; the search required the word 'challenges' in the query and title-level multimodal terms, which biases toward papers that frame problems that way; the supplementary extraction table (S1) is promised but not available, so the per-study coding cannot be checked; and there is a duplicate citation ([81] vs [74]) in Table 4. None of these sink the central taxonomy, but they are exactly the reporting gaps a referee should ask to be closed.\n\nWho this is for: researchers entering multimodal medical ML who want a quick overview of common obstacles and solution families. It is not a definitive map until the corpus attribution is cleaned up and the extraction table is released.\n\nRecommendation: worth sending to peer review, with a request that the authors fix the corpus inconsistencies, reconcile the PRISMA numbers, and make Table S1 available. The core synthesis is useful and the errors look correctable.","headline":"A useful challenge-oriented review of multimodal medical ML, but its solution map cites at least two studies outside the claimed 69-study corpus, and the PRISMA accounting needs reconciliation.","tokens_in":25887,"tokens_out":4735,"would_cite":false,"duration_ms":40937,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 69-study systematic review maps the five recurring technical obstacles in multimodal medical data modeling to the solution families researchers have proposed.","keywords":["multimodal data fusion","systematic review","missing modalities","small data","interpretability","dimensionality imbalance","fusion strategies","medical machine learning"],"falsifier":"Run the same review with the title restriction and the word 'challenges' removed, extending the search past October 2023 and including multimodal imaging, audio, and video studies; if this broader search yields additional challenge categories, changes the relative prevalence of the five, or uncovers solution families absent from the 69 studies, the review's map is incomplete. A cheaper check is to count how many of the 712 excluded records that lack the word 'challenges' in the title still report missing-modal, small-data, or interpretability methods.","tokens_in":24956,"feed_emoji":"🩺","tokens_out":4367,"duration_ms":43674,"temperature":0.7,"pith_summary":"This systematic review argues that modeling medical data from multiple modalities—imaging, genomics, electronic health records, wearables—is held back by five recurring technical obstacles: missing modalities, small datasets, interpretability, imbalance in dimensionality across modalities, and choosing the optimal fusion strategy. Across 69 included studies it collects the solution families proposed for each obstacle, such as generative imputation for missing data, transfer learning and knowledge distillation for small data, attention and gradient-based attribution for interpretability, weighted losses for dimensional imbalance, and neural architecture search for fusion design. A sympathetic reader should take the review's central contribution to be a usable challenge-to-solution map: if a practitioner faces one of these five problems, the review points to the approaches already tried and to the gaps where methods are scarce. The review also reports distributional facts about the field, including that neurology and oncology dominate the literature and that intermediate fusion is the most common strategy.","feed_headline":"Five obstacles repeat across medical multimodal AI","feed_subtitle":"A 69-study review matches each recurring challenge—missing data, small samples, black boxes—to solutions already tried.","key_machinery":"The carrying object is the challenge-to-solution taxonomy: five categories (missing modalities, small data, interpretability, imbalance in dimensionality, optimal fusion strategy) into which the 69 studies are sorted, with each study's proposed technique attached to the challenge it targets. Around this sits the four-way fusion taxonomy—early, intermediate, late, and hybrid fusion—which the review uses to characterize how studies combine modalities. The taxonomy does the work of converting a scattered literature into lookup entries: for each obstacle it names the solution families tried, the medical domains where they appear, and the combinations of modalities most often studied.","core_discovery":"The central claim is that the difficulties of multimodal medical modeling are not idiosyncratic: they fall into five named technical challenges, and the reviewed literature supplies identifiable solution families for each. Missing modalities are addressed by multitask learning, matrix completion, VAE- and GAN-based imputation, masked autoencoders, and knowledge distillation from modality-specific teachers. Small data is tackled with augmentation, transfer learning, distillation, and simulation-based knowledge transfer, though the review finds these remain resource-dependent. Interpretability is served by attention mechanisms, gradient-weighted class activation mapping, Shapley values, and biologically structured networks, but is the thinnest solution set. Dimensional imbalance is managed by weighted and focal losses, dimensionality reduction, and intermediate fusion. Optimal fusion is approached through adaptive weighting, attention, and neural architecture search over early, intermediate, late, and hybrid fusion. The review's claim is that these mappings are systematic and that the field's progress is uneven, with interpretability and truly small-data solutions the least mature.","pith_inferences":["Editorial inference: if the search had not required the word 'challenges' or a fusion/multimodal term in the title, the prevalence counts would likely shift; many papers solve these problems without framing them as challenges, so the five-way taxonomy may undercount some solution families.","Editorial inference: the review's uneven solution densities suggest a concrete research program—benchmark the missing-modality and small-data methods against each other on shared incomplete multimodal datasets, since the review catalogs options but does not rank them.","Editorial inference: dimensional imbalance and class imbalance are treated as distinct, but they likely interact; weighted and focal losses might transfer between them, since both are cases of one modality or class dominating training.","Editorial inference: the rarity of interpretability tools for multimodal models points to a testable gap—explainability methods built for single modalities may not capture cross-modal attribution, so new evaluation metrics for multimodal explanations would be needed."],"forward_implications":["A practitioner facing entirely missing modalities in a clinical dataset can go directly to the review's solution families: multitask learning per subset, matrix completion, VAE/GAN imputation, masked autoencoders, or knowledge distillation without imputation.","Teams with small datasets will find augmentation, transfer learning, and distillation, but the review implies these are partial fixes because they depend on pretrained models or large external corpora that clinical settings often lack.","Interpretability is the least-served challenge among the five; the field has fewer established tools for explaining decisions that fuse heterogeneous modalities.","Fusion design is treated as an active optimization problem: adaptive weighting and neural architecture search are the emerging answers to the question of when to fuse at data, representation, or decision level.","Imaging-plus-clinical and imaging-plus-genomic pairs dominate the 69 studies, while combinations involving wearables or three modalities are rare."],"supporting_citations":[{"why":"Establishes the motivating value of integrating electronic health records with imaging data, framing the review's scope.","marker":"[1]"},{"why":"Supplies the multimodal deep-learning fusion taxonomy and evidence that intermediate fusion handles heterogeneous and imbalanced modalities.","marker":"[3]"},{"why":"Provides criteria for choosing fusion techniques, which anchors the review's 'optimal fusion' challenge.","marker":"[5]"},{"why":"Documents missing modalities in multimodal healthcare data and motivates the challenge and its solution families.","marker":"[82]"},{"why":"Introduces the multimodal variational autoencoder that subsequent reviewed studies extend for missing-modality imputation.","marker":"[90]"},{"why":"Presents the knowledge-distillation approach for learning with incomplete modalities without imputation, a central solution family.","marker":"[70]"},{"why":"Introduces neural architecture search for multimodal fusion, the key methodological response to the optimal-fusion challenge.","marker":"[102]"},{"why":"Surveys deep multimodal learning and supplies the framing for fusion strategy selection and modality weighting.","marker":"[101]"}],"fun_headline_variants":["Five recurring medical multimodal AI problems, matched to fixes","Multimodal medical AI: five challenges, few mature fixes","The five consistent hurdles of multimodal medical machine learning","Medical multimodal AI: 5 recurring issues, patchy solutions","Five obstacles in medical multimodal modeling, solutions uneven"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the 69 studies found by its search—requiring 'model fusion', 'data fusion', or 'multimodal' in the title, the term 'challenges' in the query, and coverage only through October 2023—are representative enough of multimodal medical modeling that the five challenge categories and their solution families reflect the field's actual distribution.","fun_headline_variants_meta":{"raw":{"variants":["Five recurring medical multimodal AI problems, matched to fixes","Multimodal medical AI: five challenges, few mature fixes","The five consistent hurdles of multimodal medical machine learning","Medical multimodal AI: 5 recurring issues, patchy solutions","Five obstacles in medical multimodal modeling, solutions uneven"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1278,"prompt_tokens":866,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":334}},"tokens_in":482,"tokens_out":412,"duration_ms":4214,"temperature":1.0,"reasoning_tokens":334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:28:21.913725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same review with the title restriction and the word 'challenges' removed, extending the search past October 2023 and including multimodal imaging, audio, and video studies; if this broader search yields additional challenge categories, changes the relative prevalence of the five, or uncovers solution families absent from the 69 studies, the review's map is incomplete. A cheaper check is to count how many of the 712 excluded records that lack the word 'challenges' in the title still report missing-modal, small-data, or interpretability methods.","supporting_citations":[{"cited_title":"Multimodal generative models for scalable weakly-supervised learning","cited_arxiv_id":null,"evidence_quote":"Introduces the multimodal variational autoencoder that subsequent reviewed studies extend for missing-modality imputation."},{"cited_title":"Mufasa: Multimodal fusion architecture search for electronic health records","cited_arxiv_id":null,"evidence_quote":"Introduces neural architecture search for multimodal fusion, the key methodological response to the optimal-fusion challenge."},{"cited_title":"Deep multimodal learning: A survey on recent advances and trends","cited_arxiv_id":null,"evidence_quote":"Surveys deep multimodal learning and supplies the framing for fusion strategy selection and modality weighting."}],"review_version":1}