REVIEW 39 references
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read On a PMC-VQA subset, GRPO-based RL fine-tuning of Qwen2-VL-2B-Instruct outperforms SFT in accuracy, but the study has no error bars and several prose claims conflict with its own table.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Their main table shows the RL-tuned model reaching 58.04% accuracy, versus 52.00% for full SFT, 45.98% for LoRA SFT, and 46.97% for DPO. Removing the standard deviation normalization in the advantage calculation, a variant called Dr.GRPO, pushed accuracy to 61.09%. The paper also reports that rewarding longer reasoning chains hurts accuracy, and that adding a semantic alignment reward helps accuracy but makes the model's language less fluent according to their perplexity measure.
The findings are plausible but fragile. Accuracy numbers have no error bars, the SFT baselines are not described with enough detail, and some prose conclusions contradict the table, such as claiming better token efficiency when the model produces more thinking tokens. The comparison also lacks any direct matchup against existing medical RL-VQA systems like Med-R1.
Extended reading notes
Core claim
The abstract states: "GRPO-based RL tuning consistently outperforms standard supervised fine-tuning (SFT) in both accuracy and reasoning quality." If true, medical VQA models should be post-trained with GRPO-style RL rather than SFT alone, and reward design choices (semantic alignment, normalization removal) become first-order factors.
Load-bearing premise
The SFT baselines are trained to a fair, competitive standard. The paper cites Lee (2024) for SFT implementation but provides no SFT hyperparameters (epochs, LR, LoRA rank). The LoRA and DPO results (45.98%, 46.97%) are below the zero-shot base model (47.29%), which is unusual and suggests the SFT configurations may be suboptimal. If SFT is undertrained, the headline GRPO-vs-SFT gap is not apples-to-apples.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (5)
- GRPO hyperparameters (G=8, KL=0.04, LR=1e-6, temperature 1.0, 1500 steps) =
G=8, KL=0.04, LR=1e-6, temp=1.0, steps=1500
- Semantic alignment reward prompt and binary threshold =
Yes/No from BioGPT/BioMistral, reward 1 or 0
- ECR and CWR length reward scaling =
Not specified numerically
- PMC-VQA test subset =
7K samples
- Base model choice =
Qwen2-VL-2B / Qwen2-VL-2B-Instruct
assumptions (4)
- standard math GRPO objective and advantage estimator (Eqs. 1-2) are a valid training objective.
- domain assumption PMC-VQA ground-truth answers are correct labels for the subset used.
- domain assumption LLM judges (BioGPT, BioMistral) provide a valid clinical semantic signal.
- domain assumption Evaluation metrics (similarity, perplexity, thinking reward) capture clinical quality.
Cite this review
Pith. "Pith review of Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models." pith.science (2026). https://pith.science/paper/IGX4FC6H
@misc{pith2026250513973,
author = {Pith},
title = {Pith review of: Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/IGX4FC6H}},
note = {Machine review of arXiv:2505.13973}
}
read the original abstract
Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative Policy Optimization (GRPO). However, directly applying it to medical tasks remains challenging for achieving clinically grounded model behavior. Motivated by the need to align model response with clinical expectations, we investigate four critical dimensions that affect the effectiveness of RL-based tuning in medical visual question answering (VQA): base model initialization strategy, the role of medical semantic alignment, the impact of length-based rewards on long-chain reasoning, and the influence of bias. We conduct extensive experiments to analyze these factors for medical MLLMs, providing new insights into how models are domain-specifically fine-tuned. Additionally, our results also demonstrate that GRPO-based RL tuning consistently outperforms standard supervised fine-tuning (SFT) in both accuracy and reasoning quality.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Alisson Azzolini, Hannah Brandon, Prithvijit Chattopadhyay, Huayu Chen, Jinju Chu, Yin Cui, Jenna Diamond, Yifan Ding, Francesco Ferroni, Rama Govindaraju, and 1 others. 2025. Cosmos-reason1: From physical common sense to embodied reasoning. arXiv preprint arXiv:2503.15558
arXiv 2025
-
[4]
Liang Chen, Lei Li, Haozhe Zhao, Yifan Song, and Vinci. 2025 a . R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3. https://github.com/Deep-Agent/R1-V. Accessed: 2025-02-02
work page 2025
-
[5]
Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, and Abolfazl Razi. 2025 b . Dra-grpo: Exploring diversity-aware reward adjustment for r1-zero-like training of large language models. arXiv preprint arXiv:2505.09655
arXiv 2025
-
[6]
Kanzhi Cheng, Yantao Li, Fangzhi Xu, Jianbing Zhang, Hao Zhou, and Yang Liu. 2024. Vision-language models can self-improve reasoning via reflection. arXiv preprint arXiv:2411.00855
arXiv 2024
-
[7]
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, and 1 others. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1--53
2024
-
[8]
Matthew Chung. 2025. Training a qwen 2.5 model for medical reasoning with grpo: A tutorial and “aha!” moment. Accessed: 2025-05-16
work page 2025
Show all 39 references
-
[9]
Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. 2024. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. arXiv preprint arXiv:2411.14432
2024 arXiv
-
[10]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[11]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[12]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3
2022
-
[13]
Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. 2025. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749
2025 arXiv
-
[14]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \'e lio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Tho...
2023 arXiv
-
[15]
Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip HS Torr, Salman Khan, and Fahad Shahbaz Khan. 2025. Llm post-training: A deep dive into reasoning large language models. arXiv preprint arXiv:2502.21321
2025 arXiv
-
[16]
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. Biomistral: A collection of open-source pretrained large language models for medical domains. arXiv preprint arXiv:2402.10373
2024 arXiv
-
[17]
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xiaofeng Yang. 2025. Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models. arXiv preprint arXiv:2503.13939
2025
-
[18]
Yuwon Lee. 2024. https://github.com/2U1/Qwen2-VL-Finetune Qwen2-vl-finetune
2024
-
[19]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems, 36:34892--34916
2023
-
[20]
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. 2025. Understanding r1-zero-like training: A critical perspective. arXiv preprint arXiv:2503.20783
2025 arXiv
-
[21]
INGIN LLMS. 2025. Demystifying long chain-of-thought reason. arXiv
2025
-
[22]
Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, and Hongsheng Li. 2025. Ui-r1: Enhancing action prediction of gui agents by reinforcement learning. arXiv preprint arXiv:2503.21620
2025 arXiv
-
[23]
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6):bbac409
2022
-
[24]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728--53741
2023
-
[25]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. https://arxiv.org/abs/1707.06347 Proximal policy optimization algorithms . Preprint, arXiv:1707.06347
2017 arXiv
-
[26]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300
2024 arXiv
-
[27]
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, and 1 others. 2025 a . Vlm-r1: A stable and generalizable r1-style large vision-language model. arXiv preprint arXiv:2504.07615
2025 arXiv
-
[28]
Haozhan Shen, Zilun Zhang, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao. 2025 b . Vlm-r1: A stable and generalizable r1-style large vision-language model. https://github.com/om-ai-lab/VLM-R1. Accessed: 2025-02-15
2025
-
[29]
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, and 1 others. 2025. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599
2025 arXiv
-
[30]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. https://arxiv.org/abs/2409.12191 Qwen2-v...
2024 arXiv
-
[31]
Xiaodong Wang and Peixi Peng. 2025. Open-r1-video. https://github.com/Wang-Xiaodong1899/Open-R1-Video
2025
-
[32]
Xiaobo Xia and Run Luo. 2025. Gui-r1: A generalist r1-style vision-language action model for gui agents. arXiv preprint arXiv:2504.10458
2025 arXiv
-
[33]
Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, and Yiming Yang. 2024. Improve vision language model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198
2024 arXiv
-
[34]
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and 1 others. 2023 a . Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792
2023
-
[35]
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 b . Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415
2023 arXiv
-
[36]
Jiaxing Zhao, Xihan Wei, and Liefeng Bo. 2025. R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning. arXiv preprint arXiv:2503.05379
2025 arXiv
-
[37]
Yaowei Zheng, Junting Lu, Shenzhi Wang, Zhangchi Feng andDongdong Kuang, and Yuwen Xiong. 2025. Easyr1: An efficient, scalable, multi-modality rl training framework. https://github.com/hiyouga/EasyR1
2025
-
[39]
aha moment
Hengguang Zhou, Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. 2025 b . R1-zero's" aha moment" in visual reasoning on a 2b non-sft model. arXiv preprint arXiv:2503.05132
2025 arXiv
-
[40]
Wenhui Zhu, Xin Li, Xiwen Chen, Peijie Qiu, Vamsi Krishna Vasa, Xuanzhao Dong, Yanxi Chen, Natasha Lepore, Oana Dumitrascu, Yi Su, and 1 others. 2025. Retinalgpt: A retinal clinical preference conversational assistant powered by large vision-language models. arXiv preprint arX...
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.