REVIEW 2 major objections 1 minor 114 references
MedQwen partitions pretrained weight spectra into routed low-rank experts so a medical vision-language model approaches full fine-tuning quality with hundreds of times fewer trainable parameters while sharply reducing sequential forgetting.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 11:43 UTC pith:5WFJ2FBP
load-bearing objection We only have the ScatterPrism abstract; the supplied “full text” is a different paper (MedQwen), so the CFM-pathology claim cannot be checked. the 2 major comments →
ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Initializing each LoRA expert from a distinct, non-overlapping SVD segment of the pretrained weight, then correcting residual weight mismatch and scaling the updates so that low-rank gradients match full-rank MoE dynamics, produces a parameter-efficient medical VLM that specializes without discarding useful pretraining structure and resists both cross-dataset interference and catastrophic forgetting.
What carries the argument
Sparse spectral LoRA: non-overlapping SVD segments initialize the experts; a residual compensation matrix restores the equivalent weight at initialization; and a closed-form scaling factor s* = sqrt(3n η / r) aligns each expert’s effective gradient with that of full-rank MoE fine-tuning.
Load-bearing premise
The distinct singular-value segments remain informative after damping and scaling, so the router continues to specialize rather than collapsing to a few dominant experts or behaving like ordinary zero-initialized adapters.
What would settle it
On the Harvard-FairVLMed → PathVQA sequential protocol, measure accuracy retention after 15 epochs: if MedQwen’s drop substantially exceeds the reported ~5 percent while parameter count stays at 2.24 percent of full fine-tuning, or if zero-shot radiology accuracy falls well below 95 percent of full-FT MoE, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission is titled and abstracted as ScatterPrism (arXiv:2604.01313), claiming that the standard Conditional Flow Matching (CFM) training loss is fundamentally misleading for generative simulation and inverse problems in particle and nuclear physics: on a Jefferson Lab γp→ρ⁰p→π⁺π⁻p kinematic dataset the CFM loss plateaus while physics-informed metrics continue to improve; the authors introduce ScatterPrism, synthetic 1-D stress tests, and a multi-metric diagnostic protocol to ensure kinematic fidelity without memorization, with intended extension to EIC/HEP and other domains. The body of the manuscript supplied for review, however, is an unrelated paper (Sparse Spectral LoRA / MedQwen, arXiv:2604.01310) on SVD-structured LoRA Mixture-of-Experts for medical vision–language models. No CFM derivation, ScatterPrism architecture, Jefferson Lab results, synthetic stress-test definitions, physics-informed metrics, or multi-metric protocol appear in the provided full text.
Significance. If the abstract’s claims were substantiated by a matching manuscript, the result would be of genuine interest to the generative-modeling and subatomic-physics communities: a documented, dataset-agnostic pathology of the CFM objective, together with an efficient surrogate and a concrete multi-metric protocol, would affect how flow-matching models are trained and validated for detector simulation, unfolding, and jet modeling. Because the supplied full text does not contain those claims, equations, experiments, or code, the significance of ScatterPrism cannot be assessed from the materials under review.
major comments (2)
- Manuscript identity mismatch: the title, abstract, and paper_id (2604.01313, ScatterPrism / CFM / nuclear-physics kinematics) do not correspond to the full manuscript text, which is Sparse Spectral LoRA / MedQwen (medical VLMs, arXiv:2604.01310). Consequently there are no equations for CFM or ScatterPrism, no loss curves, no Jefferson Lab kinematic results, no synthetic 1-D stress-test definitions, no physics-informed metric definitions, and no multi-metric protocol to evaluate. The central claim that “CFM loss plateaus prematurely while physics metrics continue improving” and that this is a “dataset-agnostic pathology” cannot be checked.
- Because the load-bearing evidence (NP dataset experiments, synthetic stress tests, ScatterPrism architecture, and the proposed multi-metric diagnostic) is absent from the supplied text, the weakest assumption identified in the abstract—that the loss-versus-physics disconnect is a general CFM pathology rather than an artifact of one dataset or implementation—remains untestable. No revision of the present MedQwen manuscript can repair this; the correct ScatterPrism manuscript must be provided.
minor comments (1)
- The abstract alone is well written and the scientific motivation (EIC-relevant NP kinematics, extension to HEP jets) is clear; presentation issues in the abstract are secondary to the identity mismatch.
Circularity Check
No circularity: abstract claims are empirical observations of loss-vs-metric disconnect; supplied full text is an unrelated paper (MedQwen) containing no CFM/ScatterPrism derivations to inspect.
full rationale
The only available content for ScatterPrism (arXiv 2604.01313) is its abstract, which asserts an empirical pathology (CFM training loss plateaus while physics-informed kinematic metrics continue to improve on Jefferson Lab data and synthetic 1-D stress tests) and proposes multi-metric diagnostics plus the ScatterPrism surrogate. No equations, loss definitions, fitted parameters, uniqueness theorems, or self-citations appear in the abstract that would make any claimed result equivalent to its inputs by construction. The CACHEABLE PAPER SOURCE CONTEXT instead supplies the complete unrelated manuscript of MedQwen / Sparse Spectral LoRA (arXiv 2604.01310). That manuscript's SVD-MoE initialization, residual compensation (W_res), and scaling theorems (Theorems 1-5) are self-contained algebraic derivations from gradient expressions and router moments; they do not reduce any prediction to a fitted input or load-bearing self-citation of the target claim. Because the derivation chain of the paper under review is absent, no circular step can be exhibited. Score is therefore 0 with empty steps, consistent with the default that most papers (and abstracts) contain no circularity.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Conditional Flow Matching is a valid generative model for high-dimensional kinematic distributions arising in nuclear and particle physics.
- ad hoc to paper Physics-informed metrics (unspecified in the abstract) are more faithful indicators of kinematic fidelity than the standard CFM loss.
invented entities (1)
-
ScatterPrism
no independent evidence
read the original abstract
High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation. While Conditional Flow Matching (CFM) offers a robust acceleration approach, we demonstrate its standard training loss is fundamentally misleading. Specifically, utilizing a Jefferson Lab Nuclear Physics (NP) kinematic dataset ($\gamma p \to \rho^0 p \to \pi^+\pi^- p$), we expose that CFM loss plateaus prematurely, obscuring ongoing physical refinement. To verify this disconnect is a dataset-agnostic pathology, we introduce ScatterPrism, an efficient generative surrogate evaluated against both the NP data and synthetic stress tests modeling challenging 1D distribution topologies. Coupling these benchmarks, we establish that physics-informed metrics continue improving long after standard loss converges. Consequently, we propose a multi-metric diagnostic protocol to ensure true kinematic fidelity without data memorization. Driven by NP challenges relevant to the forthcoming Electron-Ion Collider (EIC), this unified machinery has strong potential to extend to High-Energy Physics (HEP) applications, such as jet modeling. Furthermore, the framework holds promise for broader domains requiring rigorous generative reliability, including medical imaging, astrophysics, and quantitative finance.
Reference graph
Works this paper leans on
-
[1]
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. InProceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005. 13
2005
-
[2]
Performance of chatgpt on a radiology board-style examina- tion: insights into current strengths and limitations.Radiology, 307(5):e230582, 2023
Rajesh Bhayana, Satheesh Krishna, and Robert R Bleakney. Performance of chatgpt on a radiology board-style examina- tion: insights into current strengths and limitations.Radiology, 307(5):e230582, 2023. 1
2023
-
[3]
Lora learns less and forgets less.arXiv preprint arXiv:2405.09673,
Dan Biderman, Jacob Portes, Jose Javier Gonzalez Ortiz, Man- sheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, et al. Lora learns less and forgets less.arXiv preprint arXiv:2405.09673,
-
[4]
Aofei Chang, Le Huang, Parminder Bhatia, Taha Kass-Hout, Fenglong Ma, and Cao Xiao. Medheval: Benchmarking hal- lucinations and mitigation strategies in medical large vision- language models.arXiv preprint arXiv:2503.02157, 2025. 14
Pith/arXiv arXiv 2025
-
[5]
Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao, Shu- nian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, et al. Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale.arXiv preprint arXiv:2406.19280, 2024. 2
Pith/arXiv arXiv 2024
-
[6]
Shaoxiang Chen, Zequn Jie, and Lin Ma. Llava-mole: Sparse mixture of lora experts for mitigating data conflicts in instruc- tion finetuning mllms.arXiv preprint arXiv:2401.16160, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[7]
Generating radiology reports via memory-driven transformer
Zhihong Chen, Yan Song, Tsung-Hui Chang, and Xiang Wan. Generating radiology reports via memory-driven transformer. arXiv preprint arXiv:2010.16056, 2020. 13
Pith/arXiv arXiv 2010
-
[8]
Octavius: Mitigating task interference in mllms via lora-moe
Zeren Chen, Ziqin Wang, Zhen Wang, Huayang Liu, Zhenfei Yin, Si Liu, Lu Sheng, Wanli Ouyang, Yu Qiao, and Jing Shao. Octavius: Mitigating task interference in mllms via lora-moe. arXiv preprint arXiv:2311.02684, 2023. 2
Pith/arXiv arXiv 2023
-
[9]
Jiashun Cheng, Aochuan Chen, Nuo Chen, Ziqi Gao, Yuhan Li, Jia Li, and Fugee Tsung. Revisiting lora through the lens of parameter redundancy: Spectral encoding helps.arXiv preprint arXiv:2506.16787, 2025. 2
Pith/arXiv arXiv 2025
-
[10]
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y . Wu, Zhenda Xie, Y . K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024. 19
2024
-
[11]
Preparing a col- lection of radiology examinations for distribution and retrieval
Dina Demner-Fushman, Marc D Kohli, Marc B Rosen- man, Sonya E Shooshan, Laritza Rodriguez, Sameer Antani, George R Thoma, and Clement J McDonald. Preparing a col- lection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association, 23 (2):304–310, 2016. 13, 14
2016
-
[12]
Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Jun Zhao, Wei Shen, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Xiaoran Fan, et al. Loramoe: Alleviate world knowledge forgetting in large language models via moe-style plugin.arXiv preprint arXiv:2312.09979, 2023. 2
Pith/arXiv arXiv 2023
-
[13]
Make LoRA great again: Boost- ing LoRA with adaptive singular values and mixture-of-experts optimization alignment
Chenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu, Xiaoye Qu, Wei Wei, and Yu Cheng. Make LoRA great again: Boost- ing LoRA with adaptive singular values and mixture-of-experts optimization alignment. InProceedings of the 42nd Interna- tional Conference on Machine Learning, pages 15804–15832. PMLR, 2025. 3
2025
-
[14]
Switch trans- formers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23 (120):1–39, 2022
William Fedus, Barret Zoph, and Noam Shazeer. Switch trans- formers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23 (120):1–39, 2022. 19
2022
-
[15]
Yunhao Gou, Zhili Liu, Kai Chen, Lanqing Hong, Hang Xu, Aoxue Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. Mix- ture of cluster-conditional lora experts for vision-language in- struction tuning.arXiv preprint arXiv:2312.12379, 2023. 2
Pith/arXiv arXiv 2023
-
[16]
Sara: Singular-value based adaptive low-rank adaption
Jihao Gu, Shuai Chen, Zelin Wang, Yibo Zhang, and Ping Gong. Sara: Singular-value based adaptive low-rank adaption. arXiv preprint arXiv:2408.03290, 2024. 2
Pith/arXiv arXiv 2024
-
[17]
Detecting and pre- venting hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas. Detecting and pre- venting hallucinations in large vision language models. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 18135–18143, 2024. 1
2024
-
[18]
Performance of gpt-4 with vision on text-and image-based acr diagnostic radiology in-training ex- amination questions.Radiology, 312(3):e240153, 2024
Nolan Hayden, Spencer Gilbert, Laila M Poisson, Brent Grif- fith, and Chad Klochko. Performance of gpt-4 with vision on text-and image-based acr diagnostic radiology in-training ex- amination questions.Radiology, 312(3):e240153, 2024. 1
2024
-
[19]
Upcycling large language mod- els into mixture of experts.arXiv preprint arXiv:2410.07524,
Ethan He, Abhinav Khattar, Ryan Prenger, Vijay Korthikanti, Zijie Yan, Tong Liu, Shiqing Fan, Ashwath Aithal, Mohammad Shoeybi, and Bryan Catanzaro. Upcycling large language mod- els into mixture of experts.arXiv preprint arXiv:2410.07524,
-
[20]
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Peng- tao Xie. Pathvqa: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020. 6, 13, 14
Pith/arXiv arXiv 2003
-
[21]
Alternate low-rank matrix approximation in latent semantic analysis.Scientific Programming, 2019(1):1095643, 2019
Fahrettin Horasan, Hasan Erbay, Fatih Varc ¸ın, and Emre Deniz. Alternate low-rank matrix approximation in latent semantic analysis.Scientific Programming, 2019(1):1095643, 2019. 2
2019
-
[22]
Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin. Language model compression with weighted low-rank factorization. InInternational Conference on Learning Representations, 2021. 2
2021
-
[23]
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. InInternational Confer- ence on Learning Representations, 2021. 22
2021
-
[24]
Lora: Low-rank adaptation of large language models.ICLR, 1(2):3,
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models.ICLR, 1(2):3,
-
[25]
Teng Hu, Jiangning Zhang, Ran Yi, Hongrui Huang, Yabiao Wang, and Lizhuang Ma. Sara: High-efficient diffusion model fine-tuning with progressive sparse low-rank adaptation.arXiv preprint arXiv:2409.06633, 2024. 15
Pith/arXiv arXiv 2024
-
[26]
Omnimedvqa: A new large-scale com- prehensive evaluation benchmark for medical lvlm
Yutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao, and Ping Luo. Omnimedvqa: A new large-scale com- prehensive evaluation benchmark for medical lvlm. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22170–22183, 2024. 13, 14
2024
-
[27]
Quilt-1m: One million image-text pairs for histopathology.Advances in neural infor- mation processing systems, 36:37995–38017, 2023
Wisdom Ikezogwo, Saygin Seyfioglu, Fatemeh Ghezloo, Dy- lan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro. Quilt-1m: One million image-text pairs for histopathology.Advances in neural infor- mation processing systems, 36:37995–38017, 2023. 14
2023
-
[28]
Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Sil- viana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison. InProceedings of the AAAI conference on artificial intelligence, pages 590–597, 2019. 13
2019
-
[29]
Adaptive mixtures of local experts.Neural computation, 3(1):79–87, 1991
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Ge- offrey E Hinton. Adaptive mixtures of local experts.Neural computation, 3(1):79–87, 1991. 2
1991
-
[30]
Radgraph: Extracting clinical entities and relations from radiology reports
Saahil Jain, Ashwin Agrawal, Adriel Saporta, Steven QH Truong, Du Nguyen Duong, Tan Bui, Pierre Chambon, Yuhao Zhang, Matthew P Lungren, Andrew Y Ng, et al. Radgraph: Extracting clinical entities and relations from radiology reports. arXiv preprint arXiv:2106.14463, 2021. 14
Pith/arXiv arXiv 2021
-
[31]
Mixtral of experts.arXiv preprint arXiv:2401.04088, 2024
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Flo- rian Bressand, et al. Mixtral of experts.arXiv preprint arXiv:2401.04088, 2024. 4
Pith/arXiv arXiv 2024
-
[32]
Alistair EW Johnson, Tom J Pollard, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Yifan Peng, Zhiyong Lu, Roger G Mark, Seth J Berkowitz, and Steven Horng. Mimic- cxr-jpg, a large publicly available database of labeled chest ra- diographs.arXiv preprint arXiv:1901.07042, 2019. 13, 14
Pith/arXiv arXiv 1901
-
[33]
Mimic-iv, a freely acces- sible electronic health record dataset.Scientific data, 10(1):1,
Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. Mimic-iv, a freely acces- sible electronic health record dataset.Scientific data, 10(1):1,
-
[34]
Muhammad Uzair Khattak, Shahina Kunhimon, Muzammal Naseer, Salman Khan, and Fahad Shahbaz Khan. Unimed-clip: Towards a unified image-text pretraining paradigm for diverse medical imaging modalities.arXiv preprint arXiv:2412.10372,
-
[35]
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Ve- ness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13): 3521–3526, 2017. 7
2017
-
[36]
Singular value few-shot adaptation of vision-language models.arXiv preprint arXiv:2509.03740, 2025
Taha Koleilat, Hassan Rivaz, and Yiming Xiao. Singular value few-shot adaptation of vision-language models.arXiv preprint arXiv:2509.03740, 2025. 2
Pith/arXiv arXiv 2025
-
[37]
A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018. 6, 13, 14
2018
-
[38]
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. InProceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 3045–3059, 2021. 2
2021
-
[39]
Llava-med: Training a large language-and- vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and- vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564, 2023. 1, 2, 5, 6
2023
-
[40]
Dengchun Li, Yingzi Ma, Naizheng Wang, Zhengmao Ye, Zhiyuan Cheng, Yinghao Tang, Yan Zhang, Lei Duan, Jie Zuo, Cal Yang, et al. Mixlora: Enhancing large language models fine-tuning with lora-based mixture of experts.arXiv preprint arXiv:2404.15159, 2024. 1
Pith/arXiv arXiv 2024
-
[41]
Evaluating object hallucination in large vision- language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision- language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 292–305, Singapore, 2023. Association for Computational Lin- guistics. 13
2023
-
[42]
Ensembles of low-rank expert adapters.arXiv preprint arXiv:2502.00089, 2025
Yinghao Li, Vianne Gao, Chao Zhang, and MohamadAli Torkamani. Ensembles of low-rank expert adapters.arXiv preprint arXiv:2502.00089, 2025. 2
Pith/arXiv arXiv 2025
-
[43]
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. InText summarization branches out, pages 74–81,
-
[44]
Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li, Wanggui He, Hao Jiang, Mengze Li, Xiaohui Song, et al. Healthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation.arXiv preprint arXiv:2502.09838,
-
[45]
Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments
Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments. InInternational Conference on Medical Image Com- puting and Computer-Assisted Intervention, pages 525–536. Springer, 2023. 14
2023
-
[46]
Medi- cal visual question answering: A survey.Artificial Intelligence in Medicine, 143:102611, 2023
Zhihong Lin, Donghao Zhang, Qingyi Tao, Danli Shi, Gholam- reza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. Medi- cal visual question answering: A survey.Artificial Intelligence in Medicine, 143:102611, 2023. 13
2023
-
[47]
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao- Ming Wu. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In2021 IEEE 18th international symposium on biomedical imaging (ISBI), pages 1650–1654. IEEE, 2021. 6, 13, 14
2021
-
[48]
Mitigating hallucination in large multi- modal models via robust instruction tuning
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Ya- coob, and Lijuan Wang. Mitigating hallucination in large multi- modal models via robust instruction tuning. InICLR, 2024. 1
2024
-
[49]
Application of large language models in medicine
Fenglin Liu, Hongjian Zhou, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S Chen, Yining Hua, Peilin Zhou, et al. Application of large language models in medicine. Nature Reviews Bioengineering, pages 1–20, 2025. 2
2025
-
[50]
Im- proved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Im- proved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 1
2024
-
[51]
Dora: Weight-decomposed low-rank adapta- tion
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. Dora: Weight-decomposed low-rank adapta- tion. InInternational Conference on Machine Learning, pages 32100–32121. PMLR, 2024. 2, 15
2024
-
[52]
Vil- bert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vil- bert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019. 1
2019
-
[53]
Twin-merging: Dynamic integration of modular expertise in model merging
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, and Yu Cheng. Twin-merging: Dynamic integration of modular expertise in model merging. InThe Thirty-eighth An- nual Conference on Neural Information Processing Systems,
-
[54]
Tongxu Luo, Jiahe Lei, Fangyu Lei, Weihao Liu, Shizhu He, Jun Zhao, and Kang Liu. Moelora: Contrastive learning guided mixture of experts on parameter-efficient fine-tuning for large language models.arXiv preprint arXiv:2402.12851, 2024. 2
Pith/arXiv arXiv 2024
-
[55]
Fairclip: Har- nessing fairness in vision-language learning
Yan Luo, Min Shi, Muhammad Osama Khan, Muham- mad Muneeb Afzal, Hao Huang, Shuaihang Yuan, Yu Tian, Luo Song, Ava Kouhana, Tobias Elze, et al. Fairclip: Har- nessing fairness in vision-language learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12289–12301, 2024. 14
2024
-
[56]
Biomedgpt: An open multimodal large language model for biomedicine.IEEE Journal of Biomedical and Health Informatics, 2024
Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Massimo Hong, Yushuai Wu, Mu Qiao, and Zaiqing Nie. Biomedgpt: An open multimodal large language model for biomedicine.IEEE Journal of Biomedical and Health Informatics, 2024. 2
2024
-
[57]
PiSSA: Prin- cipal singular values and singular vectors adaptation of large language models
Fanxu Meng, Zhaohui Wang, and Muhan Zhang. PiSSA: Prin- cipal singular values and singular vectors adaptation of large language models. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. 2, 3, 4
2024
-
[58]
Med-flamingo: a multimodal med- ical few-shot learner
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal med- ical few-shot learner. InMachine Learning for Health (ML4H), pages 353–367. PMLR, 2023. 1
2023
-
[59]
Multimodal large language models in medical imaging: Current state and future directions
Yoojin Nam, Dong Yeong Kim, Sunggu Kyung, Jinyoung Seo, Jeong Min Song, Jimin Kwon, Jihyun Kim, Wooyoung Jo, Hyungbin Park, Jimin Sung, et al. Multimodal large language models in medical imaging: Current state and future directions. Korean Journal of Radiology, 26(10):900, 2025. 2
2025
-
[60]
D-rax: Domain-specific radiologic assistant leveraging multi-modal data and expert model predictions
Hareem Nisar, Syed Muhammad Anwar, Zhifan Jiang, Abhi- jeet Parida, Ramon Sanchez-Jacob, Vishwesh Nath, Holger R Roth, and Marius George Linguraru. D-rax: Domain-specific radiologic assistant leveraging multi-modal data and expert model predictions. InInternational Workshop on Foundation Models for General Medical AI, pages 91–102. Springer, 2024. 1
2024
-
[61]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318,
-
[62]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pages 8748–8763. PMLR,
-
[63]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538, 2017. 2
Pith/arXiv arXiv 2017
-
[64]
Toward expert-level med- ical question answering with large language models.Nature Medicine, 31(3):943–950, 2025
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. Toward expert-level med- ical question answering with large language models.Nature Medicine, 31(3):943–950, 2025. 2
2025
-
[65]
Combining automatic labelers and expert annotations for accurate radiology report labeling using bert
Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, An- drew Y Ng, and Matthew Lungren. Combining automatic labelers and expert annotations for accurate radiology report labeling using bert. InProceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), pages 1500–1519, 2020. 13, 14
2020
-
[66]
Dimen- sionality reduction using pca and svd in big data: A compar- ative case study
Sudeep Tanwar, Tilak Ramani, and Sudhanshu Tyagi. Dimen- sionality reduction using pca and svd in big data: A compar- ative case study. InFuture Internet Technologies and Trends: First International Conference, ICFITT 2017, Surat, India, Au- gust 31-September 2, 2017, Proceedings 1, pages 116–125. Springer, 2018. 2
2017
-
[67]
Xraygpt: Chest radiographs summarization using large med- ical vision-language models
Omkar Chakradhar Thawakar, Abdelrahman M Shaker, Sa- hal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Khan. Xraygpt: Chest radiographs summarization using large med- ical vision-language models. InProceedings of the 23rd work- shop on biomedical natural language processing, pages 440– 448, 2024. 2
2024
-
[68]
Hydralora: An asymmetric lora architecture for efficient fine-tuning, 2024
Chunlin Tian, Zhan Shi, Zhijiang Guo, Li Li, and Chengzhong Xu. Hydralora: An asymmetric lora architecture for efficient fine-tuning, 2024. 3, 4, 5
2024
-
[69]
Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.Nature Biomedical Engineering, 6(12): 1399–1406, 2022
Ekin Tiu, Ellie Talius, Pujan Patel, Curtis P Langlotz, An- drew Y Ng, and Pranav Rajpurkar. Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.Nature Biomedical Engineering, 6(12): 1399–1406, 2022. 1
2022
-
[70]
Vigc: Visual instruction generation and correction
Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, et al. Vigc: Visual instruction generation and correction. InProceed- ings of the AAAI Conference on Artificial Intelligence, pages 5309–5317, 2024. 1
2024
-
[71]
Kasa: Knowledge-aware singular-value adaptation of large language models, 2024
Fan Wang, Juyong Jiang, Chansung Park, Sunghun Kim, and Jing Tang. Kasa: Knowledge-aware singular-value adaptation of large language models, 2024. 2, 3
2024
-
[72]
Milora: Harnessing minor singular components for parameter-efficient llm finetuning, 2024
Hanqing Wang, Yixia Li, Shuo Wang, Guanhua Chen, and Yun Chen. Milora: Harnessing minor singular components for parameter-efficient llm finetuning, 2024. 2
2024
-
[73]
Roselora: Row and column-wise sparse low-rank adaptation of pre-trained language model for knowl- edge editing and fine-tuning
Haoyu Wang, Tianci Liu, Ruirui Li, Monica Xiao Cheng, Tuo Zhao, and Jing Gao. Roselora: Row and column-wise sparse low-rank adaptation of pre-trained language model for knowl- edge editing and fine-tuning. InProceedings of the 2024 Con- ference on Empirical Methods in Natural Language Process- ing, pages 996–1008, 2024. 2
2024
-
[74]
Hanqing Wang, Zeguan Xiao, Yixia Li, Shuo Wang, Guanhua Chen, and Yun Chen. Milora: Harnessing minor singular com- ponents for parameter-efficient llm finetuning.arXiv preprint arXiv:2406.09044, 2024. 2
Pith/arXiv arXiv 2024
-
[75]
Milora: Harnessing minor singular components for parameter-efficient llm finetuning
Hanqing Wang, Yixia Li, Shuo Wang, Guanhua Chen, and Yun Chen. Milora: Harnessing minor singular components for parameter-efficient llm finetuning. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Asso- ciation for Computational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers), pages 4823–4836, 2025. 2, 3
2025
-
[76]
Lora-ga: Low-rank adaptation with gradient approximation.Advances in Neural Information Processing Systems, 37:54905–54931, 2024
Shaowen Wang, Linxi Yu, and Jian Li. Lora-ga: Low-rank adaptation with gradient approximation.Advances in Neural Information Processing Systems, 37:54905–54931, 2024. 2
2024
-
[77]
Lora-ga: Low-rank adaptation with gradient approximation
Shaowen Wang, Linxi Yu, and Jian Li. Lora-ga: Low-rank adaptation with gradient approximation. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems,
-
[78]
Xin Wang, Yu Zheng, Zhongwei Wan, and Mi Zhang. Svd-llm: Truncation-aware singular value decomposition for large lan- guage model compression.arXiv preprint arXiv:2403.07378,
-
[79]
Medclip: Contrastive learning from unpaired medical im- ages and text
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. Medclip: Contrastive learning from unpaired medical im- ages and text. InProceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Em- pirical Methods in Natural Language Processing, page 3876,
-
[80]
Lora-pro: Are low-rank adapters properly optimized?,
Zhengbo Wang, Jian Liang, Ran He, Zilei Wang, and Tieniu Tan. Lora-pro: Are low-rank adapters properly optimized?,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.