Pith. sign in

REVIEW 3 major objections 7 minor 70 references

UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment

T0 review · 3 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Requiring each modality to predict alone removes the text shortcut and improves multi-modal diagnosis.

desk verdict UniMod is a solid empirical paper with honest diagnostics, but its central story (independent supervision removes shortcuts) is not properly isolated: the missing IFE ablation and a contradictory stray number keep it from being fully convincing. read the letter →

arxiv 2608.10316 v1 pith:CNNIKPCS submitted 2026-08-10 cs.CV cs.LGcs.MM

classification cs.CVcs.LGcs.MM
keywords multi-modallearningshortcutmedicaldiagnosisvision-languagemodelsrepresentationalignmentsupervisedcontrastiveglaucomadetectionchestX-rayclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

UniMod addresses a failure mode in multi-modal medical diagnosis: when a model is trained on images plus clinical text, it often solves the task almost entirely from text, because notes contain explicit diagnostic phrases while the image patterns are harder to learn. The paper's central claim is that this shortcut disappears when the training objective forces each modality to make its own correct prediction, in addition to the fused prediction. On two benchmarks this yields substantially better AUC than gradient-based modality-balancing methods, and it keeps the model accurate when text or images are missing at test time. The broader point is that shortcut learning is not a gradient-balance problem but an objective-design problem.

What carries the argument

The load-bearing mechanism is the modality-separated attention mask combined with independent unimodal classification heads. The mask allows image tokens to attend only to image tokens and text tokens only to text tokens, so the mean-pooled image and text embeddings are genuinely unimodal; the final prediction token is the only place cross-modal interaction occurs. Three classification losses—on the image-only, text-only, and multi-modal paths—then force each modality to extract diagnostic features on its own, while an MSE cross-modal alignment loss transfers knowledge between modalities and a supervised contrastive loss structures each modality's representation space by diagnosis. The framework is assembled from established pieces, but the paper's claim is that the independent supervision, not gradient modulation, is what removes the shortcut.

What would settle it

Take the cleaned Harvard-Glaucoma and CheXpert Plus notes and run an independent audit for diagnostic cues, then train the text-only branch on examples that contain no phrase a keyword classifier would flag; if text-only AUC collapses or if a held-out keyword classifier can predict the label from the cleaned text at high accuracy, the shortcut-learning story and the comparison to gradient balancing would need re-quantification.

Watch

Extended reading notes

Core claim

The paper claims that supervising image-only, text-only, and multi-modal predictions at the same time removes the lowest-loss route to shortcut learning: with the fused loss alone, a model can satisfy the objective by riding the easier text branch while the image encoder stays nearly non-diagnostic; with independent unimodal losses, neither branch can defer to the other. UniMod operationalizes this with a custom attention mask that keeps image and text token streams separate except at a final prediction token, mean-pooled unimodal embeddings, cross-modality alignment that pulls same-patient image and text representations together, and supervised contrastive alignment that clusters same-diagnosis patients within each modality. The reported result is 0.850 AUC on Harvard-Glaucoma and 0.966 AUC on CheXpert Plus, outperforming OGM-GE and Gradient Blending by 1.6-1.8% and over 5% respectively, and a 5-class multi-label extension improves mean AUC by 0.097 over CGGM without architectural change.

Load-bearing premise

The load-bearing premise is that the text-cleaning step removes all diagnostic label leakage from the clinical notes, so the text-only supervision is learning from genuine clinical reasoning rather than from leaked keywords; if leakage remains, the shortcut is not actually closed.

Editorial extensions

If this is right

  • If the central claim is right, gradient-balancing methods such as OGM-GE and Gradient Blending address the symptom rather than the cause; changing the training objective to include unimodal supervision is the effective intervention.
  • A model trained with UniMod retains much of its accuracy when one modality is absent at test time, which matters clinically because notes may be incomplete or imaging-only diagnosis may be required.
  • The image branch becomes genuinely diagnostic: on CheXpert Plus the image-text AUC gap shrinks from 0.58 under standard multimodal LoRA to 0.06 under UniMod.
  • The method transfers to multi-label diagnosis with no architectural change, suggesting the objective, not task-specific design, drives the gain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: any vision-language model with shared or separate encoders should benefit from replacing the fused-only objective with per-modality supervision whenever one modality is cheaper to exploit than the other.
  • Beyond the paper: the attention mask changes accuracy little, but the paper argues it is a correctness precondition; an external reader could verify by checking whether the text branch's learned features change when the mask is removed.
  • Beyond the paper: a direct extension would audit the cleaned notes for residual leaked diagnostic cues; if leakage survives cleaning, the text-only baseline's high AUC could be inflated, which would change how the shortcut is quantified.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper proposes UniMod, a multi-modal medical diagnosis framework that combines fundus images and clinical notes (Harvard-Glaucoma) or chest X-rays and radiology reports (CheXpert Plus). The method adds three classification losses on image-only, text-only, and multi-modal predictions, enforces a modality-separated attention mask, and adds cross-modality MSE alignment plus within-modality supervised contrastive learning, with GradNorm weighting and LoRA fine-tuning of an InternVL2.5-8B backbone. The authors report AUC improvements over OGM-GE and Gradient Blending (0.850 vs 0.835/0.837 on Harvard-Glaucoma; 0.966 vs 0.918/0.919 on CheXpert Plus), missing-modality robustness, and a 5-class multi-label extension that improves mean AUC over CGGM by 0.097. The paper also provides loss-level diagnostics intended to show that standard multi-modal training satisfies the fused objective while leaving the image branch under-optimized.

Significance. If the central causal claim holds, the paper makes a useful and clean contribution: it identifies that gradient-level balancing does not alter the training objective, and that directly supervising each modality's independent prediction is a simple mechanism for discouraging shortcut learning in vision-language medical models. The modality-separated attention design is justified as a correctness precondition rather than an accuracy device, and the missing-modality robustness results in Table 4 and Figures 5-6 are compelling evidence that UniMod produces more balanced modality reliance. The extension to multi-label diagnosis without architectural change is also a positive feature. However, the empirical case for the core mechanism is weakened by the absence of a direct ablation of independent feature extraction, by an internal inconsistency between the reported full-model AUC values, and by incomplete documentation of the label-leakage cleaning step. The paper's significance therefore depends on whether these points can be resolved in revision.

major comments (3)
  1. [Section 5.4, Table 3] The central mechanism of the paper, Independent Feature Extraction (IFE), is never directly ablated. Table 3 removes only Within and Cross, and both rows keep IFE enabled. No row removes L_img_cls and L_txt_cls while retaining L_cross and L_within, even though the contribution list claims this is the decisive component. The only quantitative evidence for IFE is the statement that 'replacing it under the same alignment losses reaches 0.849 AUC against 0.857 for the full model'; this result has no experimental detail, no variance, and conflicts with Table 1, where UniMod achieves 0.850 AUC on Harvard-Glaucoma. Please add a w/o IFE ablation row with the same alignment losses and report the full model's AUC consistently, together with seed-level variance.
  2. [Section 5.1, Addressing potential label leakage] The text-cleaning description is not sufficient to establish that clinical notes are free of label leakage. The claim that '0% of samples contain direct label leakage' after cleaning needs the complete cleaning pattern list and an audit procedure; otherwise the text-only baseline and UniMod's text branch could still exploit diagnostic keywords, which would undermine the shortcut-learning narrative and the comparison with gradient-balancing methods. Please provide the full list of removed patterns, representative cleaned and uncleaned examples, and a quantitative check such as text-only AUC on raw versus cleaned notes.
  3. [Section 5.2, Table 1] The main results in Table 1 are reported on a single split without error bars, even though Table 5 reports means over three seeds. On Harvard-Glaucoma the reported improvement over Gradient Blending is 0.850 versus 0.837 AUC, a 0.013 difference that could plausibly lie within seed-to-seed variation for this setup. Please report mean and standard deviation over at least three seeds for the main results and state whether the differences against OGM-GE and G-Blend are statistically significant.
minor comments (7)
  1. [Abstract] There is a missing space in 'We proposeUniMod'; similar spacing issues with 'UniMod' appear in the body text.
  2. [Section 5.1, Table 2] Table references alternate between 'Table' and 'Tbl.'; please use a consistent style.
  3. [Table 1, Zero-shot row] The recall of 1.000 with AUC 0.472 on Harvard-Glaucoma suggests the zero-shot model is effectively predicting all samples as positive; please clarify the thresholding procedure and discuss the below-chance AUC.
  4. [Section 5.4, Table 3] The text mentions GradNorm and modality-separation ablations with specific AUC deltas, but Table 3 does not include these rows. Please either add them to the ablation table or clearly state that they are reported only in the text.
  5. [Section 4.1, 'Why separate the streams'] The discussion of the mask ablation in Section 4.1 would fit more naturally in Section 5.4, and the 'seed-to-seed standard deviation' is mentioned without reporting the actual standard deviation values.
  6. [Figure 3] The 'greener is better optimized' convention is not accessible to color-blind readers; please add numeric loss values or use a non-color visual cue.
  7. [Section 5.1, default cleaned text] Please state explicitly whether the text-cleaning is also applied to the text-only baseline and to the case-study reports shown in Figure 4.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: UniMod's unimodal-supervision mechanism is an intervention tested on external benchmarks; self-citations are peripheral.

full rationale

The paper's central claim is that adding image-only and text-only classification losses to the fused objective removes the lowest-loss shortcut and forces each modality to become diagnostic. That claim is an architectural intervention with the resulting AUCs evaluated on held-out test splits against externally published baselines (OGM-GE, Gradient Blending, CGGM), not a quantity derived from its own fitted constants. The unimodal losses L_img_cls and L_txt_cls are defined on separate representations and are not, by construction, equal to the multi-modal loss or to the reported AUC improvements. The paper's supporting mechanism evidence (CE_mm, CE_img, CE_txt and missing-modality robustness) is observational and could be debated, but that is an experimental-support or correctness concern, not circularity. The self-citations to Kan et al., Zheng et al., and MuteBench ([28], [29], [70]) appear only in related-work context and are not load-bearing: no uniqueness theorem, ansatz, or fitted result is imported from prior work by the same authors. The skeptical concern about the missing IFE-only ablation (Table 3 removes only Within and Cross) is a gap in isolating the independent feature extraction effect, not a reduction of the prediction to the input. Because the method is evaluated against external benchmarks with no fitted parameter renamed as a prediction and no self-citation chain supporting the main mechanism, the appropriate circularity score is 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The method relies on standard pretrained VLM representations and a hand-chosen SupCon temperature. No new entities are introduced; the core assumptions are the meaningfulness of direct MSE alignment in the shared space and the effectiveness of text cleaning in removing label leakage.

free parameters (1)
  • SupCon temperature tau = 0.07
    Chosen by hand as a standard value for supervised contrastive learning; not fitted to data, but a hand-picked hyperparameter that affects the within-modality alignment strength.
assumptions (3)
  • domain assumption InternVL2.5-8B's pre-trained weights place image and text tokens in a common embedding space where direct MSE alignment (Eq. 7) is a useful knowledge-transfer signal without a learned projection.
    Section 4.2 applies L_cross directly to z_img and z_txt, relying on the VLM's shared space. The paper does not analyze the geometry of these representations or compare to a learned projection.
  • domain assumption The text cleaning in Section 5.1 removes every direct diagnostic cue, leaving 0% of samples with direct label leakage.
    This is the foundation for the shortcut-learning measurement; the full cleaning pattern set is not listed in the paper, so an independent check is not possible.
  • standard math GradNorm provides a stable, near-optimal adaptive weighting of the classification and alignment losses.
    Section 4.4 adopts GradNorm from Chen et al. [8] as a standard multi-task learning tool; the paper does not compare to a fixed-weight sweep.

how reviews work

0 comments
Cite this review

Pith. "Pith review of UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment." pith.science (2026). https://pith.science/paper/CNNIKPCS

@misc{pith2026260810316,
  author       = {Pith},
  title        = {Pith review of: UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CNNIKPCS}},
  note         = {Machine review of arXiv:2608.10316}
}
read the original abstract

Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). We propose UniMod, a framework that mitigates shortcut learning by requiring each modality to predict the diagnosis on its own. It supervises image-only, text-only, and multi-modal classification simultaneously, so each modality must extract diagnostic features. We add cross-modality alignment for knowledge transfer and within-modality supervised contrastive alignment over same-diagnosis patients. On Harvard-Glaucoma, UniMod reaches 0.850 AUC, outperforming OGM-GE and Gradient Blending by 1.6-1.8%; on CheXpert Plus, it reaches 0.966 AUC, surpassing them by over 5%. UniMod also extends to 5-class multi-label diagnosis without architectural change, improving mean AUC by 0.097 over CGGM.

Figures

Figures reproduced from arXiv: 2608.10316 by the authors.

Figure 1
Figure 1. Comparison of Standard Multimodal LoRA and [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of UniMod. We implement a custom attention mask to enforce modality isolation, yielding a unimodal embedding for each modality. We apply three losses: (1) independent classification losses on image-only, text-only, and multi-modal predictions; (2) cross-modality alignment (Lcross) for cross-modality knowledge transfer; and (3) within-modality contrastive alignment (Lwithin) for class-level clustering. Fundu… view at source ↗
Figure 3
Figure 3. Classification loss components at the end of training [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 7
Figure 7. Figure 7: GradNorm weight evolution. 𝜆1 is fixed at 1.0; 𝜆2 (cross￾modal) and 𝜆3 (within-modal) are adjusted automatically, rising up to fourfold on CheXpert Plus, consistent with its greater modality imbalance. Harvard-Glaucoma case the note is suggestive but inconclusive (e.g.…
Figure 6
Figure 6. Figure 6: Modality-specific AUC on CheXpert Plus; annotations [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 46 canonical work pages

  1. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...

  2. [2]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2. 5-vl technical re...

  3. [3]

    Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multi- modal Machine Learning: A Survey and Taxonomy.IEEE Transactions on Pattern Analysis and Machine Intelligence(2018)

  4. [4]

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arber, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258(2021)

  5. [5]

    Langlotz

    Pierre Chambon, Jean-Benoit Delbrouck, Thomas Sounack, Shih-Cheng Huang, Zhihong Chen, Maya Varma, Steven QH Truong, Chu The Chuong, and Curtis P. Langlotz. 2024. CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats. arXiv preprint arXiv:2405.19538(2024)

  6. [6]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. InInterna- tional conference on machine learning. PmLR, 1597–1607

  7. [7]

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. InEuropean Conference on Computer Vision. 104–120

  8. [8]

    Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. 2018. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. InInternational conference on machine learning. PMLR, 794–803

Show all 70 references
  1. [9]

    Zhihong Chen, Yuhao Du, Jinpeng Hu, Yang Liu, Guanbin Li, Xiang Wan, and Tsung-Hui Chang. 2022. Multi-modal masked autoencoders for medical vision- and-language pre-training. InInternational Conference on Medical Image Comput- ing and Computer-Assisted Intervention. 679–689

  2. [10]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271(2024)

  3. [11]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic...

  4. [12]

    Alex J DeGrave, Joseph D Janizek, and Su-In Lee. 2021. AI for radiographic COVID-19 detection selects shortcuts over signal.Nature Machine Intelligence3, 7 (2021), 610–619

  5. [13]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs.Advances in Neural Information Processing Systems36 (2023), 10088–10115

  6. [14]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InInternational...

  7. [15]

    Tom Fawcett. 2006. An introduction to ROC analysis.Pattern Recognition Letters 27, 8 (2006), 861–874

  8. [16]

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learn- ing in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673

  9. [17]

    Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhao- han Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko

    Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhao- han Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap your own...

  10. [18]

    Zirun Guo, Tao Jin, and Zhou Zhao. 2024. Classifier-Guided Gradient Modulation for Enhanced Multimodal Learning. InAdvances in Neural Information Processing Systems

  11. [19]

    Haibo He and Edwardo A Garcia. 2009. Learning from imbalanced data.IEEE Transactions on Knowledge and Data Engineering21, 9 (2009), 1263–1284

  12. [20]

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum Contrast for Unsupervised Visual Representation Learning. InIEEE/CVF Conference on Computer Vision and Pattern Recognition

  13. [21]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. InInternational Conference on Machine Learning. 2790–2799

  14. [22]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations

  15. [23]

    Shih-Cheng Huang, Liyue Shen, Matthew P Lungren, and Serena Yeung. 2021. Gloria: A multimodal global-local representation learning framework for label- efficient medical image recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3942–3951

  16. [24]

    Fushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang, and Song Guo. 2024. C2KD: Bridging the Modality Gap for Cross-Modal Knowledge Distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

  17. [25]

    Mong, Safwan S

    Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, Jayne Seekins, David A. Mong, Safwan S. Halabi, Jesse K. Sandberg, Ricky Jones, David B. Larson, Curtis P. Langlotz, Bhavik ...

  18. [26]

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision- language representation learning with noisy text supervision. InInternational Conference on Machine Learning. 4904–4916

  19. [27]

    Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih ying Deng, Roger G Mark, and Steven Horng. 2019. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports.Scientific Data6, 1 (2019), 317

  20. [28]

    Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, and Tianlong Chen. 2026. TRACE: A Temporal Conditional Estimation for Multimodal Time Series Foundation Models.arXiv preprint arXiv:2606.06285(2026)

  21. [29]

    Ziwen Kan, Wugeng Zheng, Tianlong Chen, and Song Wang. 2026. PAMF: Prior-Aware Multimodal Fusion for Incomplete Time Series Data.arXiv preprint arXiv:2606.06328(2026)

  22. [30]

    Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7482–7491

  23. [31]

    Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised Contrastive Learning. InAdvances in Neural Information Processing Systems

  24. [32]

    Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language trans- former without convolution or region supervision. InInternational Conference on Machine Learning. 5583–5594

  25. [33]

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023. Llava-med: Train- ing a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems36 (20...

  26. [34]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.arXiv preprint arXiv:2301.12597(2023)

  27. [35]

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language repre- sentation learning with momentum distillation.Advances in Neural Information Processing Systems34 (2021), 9694–9705

  28. [36]

    Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. InProceedings of the 59th Annual Meeting of the Associa- tion for Computational Linguistics. 4582–4597. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Gu et al

  29. [37]

    Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022. Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning.Advances in Neural Information Processing Systems(2022)

  30. [38]

    Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Pmc-clip: Contrastive language-image pre-training using biomedical documents.arXiv preprint arXiv:2303.07240(2023)

  31. [39]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruc- tion tuning. InAdvances in Neural Information Processing Systems

  32. [40]

    Xiaoxuan Liu, Livia Faes, Aditya U Kale, Siegfried K Wagner, Dun Jack Fu, Alice Bruynseels, Thushika Mahendiran, Gabriella Moraes, Mohith Shamdas, and Christoph Kern. 2019. A comparison of deep learning performance against health- care professionals in detecting diseases from ...

  33. [41]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of the IEEE/CVF International Conference on Computer Vision. 10012–10022

  34. [42]

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretrain- ing task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems32 (2019)

  35. [43]

    Yan Luo, Min Shi, Yu Tian, Tobias Elze, and Mengyu Wang. 2023. Har- vard glaucoma detection and progression: A multimodal multitask dataset and generalization-reinforced semi-supervised learning. InProceedings of the IEEE/CVF International Conference on Computer Vision. 20471–20482

  36. [44]

    Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout. 2016. ModDrop: Adaptive multi-modal gesture recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence38, 8 (2016), 1692–1706

  37. [45]

    Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding.arXiv preprint arXiv:1807.03748(2018)

  38. [46]

    OpenAI. 2023. GPT-4V(ision) System Card. https://openai.com/research/gpt-4v- system-card

  39. [47]

    Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8238–8247

  40. [48]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual mod- els from natural language supervision. InInternatio...

  41. [49]

    Lungren, and Andrew Y

    Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu, Brandon Yang, Hershel Mehta, Tony Duan, Daisy Ding, Aarti Bagul, Curtis Langlotz, Katie Shpanskaya, Matthew P. Lungren, and Andrew Y. Ng. 2017. CheXNet: Radiologist-level pneumonia detec- tion on chest x-rays with deep learning.arXiv...

  42. [50]

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. InInternational Conference on Medical image computing and computer-assisted intervention. Springer, 234–241

  43. [51]

    Sebastian Ruder. 2017. An overview of multi-task learning in deep neural net- works.arXiv preprint arXiv:1706.05098(2017)

  44. [52]

    Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2020. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. InInternational Conference on Learning Representations

  45. [53]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Ague...

  46. [54]

    Yih-Chung Tham, Xiang Li, Tien Y Wong, Harry A Quigley, Tin Aung, and Ching- Yu Cheng. 2014. Global prevalence of glaucoma and projections of glaucoma burden through 2040: a systematic review and meta-analysis.Ophthalmology 121, 11 (2014)

  47. [55]

    Ekin Tiu, Ellie Talius, Pujan Patel, Curtis P Langlotz, Andrew Y Ng, and Pranav Rajpurkar. 2022. Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning.Nature Biomedical Engineering6, 12 (2022), 1399–1406

  48. [56]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. Llama: Open and efficient foundation ...

  49. [57]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Processing Systems. 5998–6008

  50. [58]

    Tongzhou Wang and Phillip Isola. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. InInternational Conference on Machine Learning

  51. [59]

    Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. Internvl3. 5: Ad- vancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265(2025)

  52. [60]

    Weiyao Wang, Du Tran, and Matt Feiszli. 2020. What makes training multi- modal classification networks hard?. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12695–12705

  53. [61]

    Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. 2022. Medclip: Contrastive learning from unpaired medical images and text.arXiv preprint arXiv:2210.10163(2022)

  54. [62]

    Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Medklip: Medical knowledge enhanced language-image pre-training.medRxiv (2023), 2023–01

  55. [63]

    Peng Xu, Xiatian Zhu, and David A Clifton. 2023. Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 10 (2023), 12113–12132

  56. [64]

    Jinhui Yi, Huan Yan, Haotian Wang, Jian Yuan, and Yong Li. 2023. Deepsta: A spatial-temporal attention network for logistics delivery timely rate prediction in anomaly conditions. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management. 4916–4922

  57. [65]

    Jinhui Yi, Huan Yan, Haotian Wang, Jian Yuan, and Yong Li. 2024. Learning to Estimate Package Delivery Time in Mixed Imbalanced Delivery and Pickup Logistics Services. InProceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems. 432–443

  58. [66]

    Kihyun You, Jawook Gu, Jiyeon Ham, Beomhee Park, Jiho Kim, Eun K Hong, Woonhyuk Baek, and Byungseok Roh. 2023. Cxr-clip: Toward large scale chest x-ray language-image pre-training. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention. 101–111

  59. [67]

    Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021. Bar- low twins: Self-supervised learning via redundancy reduction. InInternational Conference on Machine Learning. 12310–12320

  60. [68]

    Lungren, Tristan Naumann, Sheng Wang, and Hoifung Poon

    Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Andrea Tupini, Yu Wang, Matt Mazzola, Swadheen Shukla, Lars Liden, Jianfeng Gao, Angela Crabtree, Brian Piening, Carlo Bifulco, Matthew P....

  61. [69]

    Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu. 2018. Deep mutual learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4320–4328

  62. [70]

    Wugeng Zheng, Ziwen Kan, Tianlong Chen, Chen Chen, and Song Wang. 2026. MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multi- modal Fusion.arXiv preprint arXiv:2605.15235(2026)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.