REVIEW 3 cited by
The Hallucination Tax of Reinforcement Finetuning
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Standard RFT sharply reduces LLM refusal on unanswerable questions, and adding 10% synthetic unanswerable math during RFT restores refusal with small accuracy losses.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
To fix the problem, they build a dataset called SUM. They take competition math problems and modify them with five recipes: delete key information, make information ambiguous, add impossible conditions, mention objects that are absent from the problem, or delete the question itself. During reinforcement finetuning, they replace 10% of the training data with these unanswerable variants. The reward function already gives credit for correct answers; they extend it so that refusing an unanswerable problem with a specific "I don't know" phrase also earns credit. After training, models refuse more often on the SUM test set, on human-written unanswerable math problems, and on factual unanswerable questions, while math accuracy on solvable problems drops only a few points.
The main caveat is that models are literally rewarded for refusing on exactly the kind of data used to test them, and the paper does not measure whether they start refusing too many answerable factual questions. So the stronger claim that models genuinely reason about their own uncertainty is not fully proven, but the training recipe itself is simple and likely useful.
Extended reading notes
Core claim
The paper claims that standard reinforcement finetuning (RFT) degrades refusal behavior: "standard RFT training could reduce model refusal rates by more than 80%, which significantly increases model's tendency to hallucinate," and that "incorporating just 10% SUM during RFT substantially restores appropriate refusal behavior, with minimal accuracy trade-offs on solvable tasks." If correct, RFT for math improves accuracy at the cost of epistemic humility, and a simple 10% synthetic unanswerable-data mix recovers most of that humility while preserving most accuracy, including transfer to out-of-domain math and to factual unanswerable questions.
Load-bearing premise
The paper's strongest generalization claim rests on the unstated premise that higher refusal rates on unanswerable benchmarks reflect genuine uncertainty reasoning rather than a generic over-refusal tendency. The evaluation in Table 2 measures refusal on unanswerable sets but never measures false refusals on answerable factual questions, and the paper does not analyze chain-of-thought traces to show the model actually detected missing information. If the model simply learned a superficial cue to refuse, the SelfAware gains would come with unacceptable over-refusal in real applications. This premise enters in Section 5.2, Finding 2, where the authors interpret the SelfAware refusal gain as evidence that models "leverage inference-time compute to reason about their own uncertainty."
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (1)
- SUM mixing ratio =
10% (with 0%, 1%, 30%, 50% also evaluated)
assumptions (4)
- domain assumption Unanswerable-problem refusal is a valid proxy for hallucination risk in deployed LLMs.
- domain assumption o3-mini generated modifications preserve genuine unanswerability at the reported correctness rate.
- ad hoc to paper The reward function in Eq. (3) correctly balances correctness and abstention.
- domain assumption The labels in UMWP and SelfAware are reliable ground truth for unanswerability.
invented entities (2)
-
Hallucination tax (coined construct)
-
Synthetic Unanswerable Math (SUM) dataset
independent evidence
Cite this review
Pith. "Pith review of The Hallucination Tax of Reinforcement Finetuning." pith.science (2026). https://pith.science/paper/CUMTZLPD
@misc{pith2026250513988,
author = {Pith},
title = {Pith review of: The Hallucination Tax of Reinforcement Finetuning},
year = {2026},
howpublished = {\url{https://pith.science/paper/CUMTZLPD}},
note = {Machine review of arXiv:2505.13988}
}
read the original abstract
Reinforcement finetuning (RFT) has become a standard approach for enhancing the reasoning capabilities of large language models (LLMs). However, its impact on model trustworthiness remains underexplored. In this work, we identify and systematically study a critical side effect of RFT, which we term the hallucination tax: a degradation in refusal behavior causing models to produce hallucinated answers to unanswerable questions confidently. To investigate this, we introduce SUM (Synthetic Unanswerable Math), a high-quality dataset of unanswerable math problems designed to probe models' ability to recognize an unanswerable question by reasoning from the insufficient or ambiguous information. Our results show that standard RFT training could reduce model refusal rates by more than 80%, which significantly increases model's tendency to hallucinate. We further demonstrate that incorporating just 10% SUM during RFT substantially restores appropriate refusal behavior, with minimal accuracy trade-offs on solvable tasks. Crucially, this approach enables LLMs to leverage inference-time compute to reason about their own uncertainty and knowledge boundaries, improving generalization not only to out-of-domain math problems but also to factual question answering tasks.
Figures
Forward citations
Cited by 3 Pith papers
-
Mechanistic Attention Guidance for Agent Memory Refinement
Attention patterns reveal how an AI agent uses memory, and using these patterns to rewrite memory improves task performance and memory efficiency.
-
Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration
A training method using reinforcement learning and answerability heuristics lets small language models actively ask for missing math details and then solve problems, raising accuracy on the new GSM-MC benchmark from 0...
-
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, and 4 others. 2025. https://arxi...
arXiv 2025
-
[4]
Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, and Pascale Fung. 2025. https://arxiv.org/abs/2504.17550 Hallulens: Llm hallucination benchmark . Preprint, arXiv:2504.17550
arXiv 2025
-
[5]
Forrest Bao, Miaoran Li, Rogger Luo, and Ofer Mendelevitch. 2024. https://doi.org/10.57967/hf/3240 HHEM-2.1-Open
doi:10.57967/hf/3240 2024
-
[6]
Assaf Ben-Kish, Moran Yanuka, Morris Alper, Raja Giryes, and Hadar Averbuch-Elor. 2023. Mitigating open-vocabulary caption hallucinations. arXiv preprint arXiv:2312.03631
arXiv 2023
-
[7]
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, and Xipeng Qiu. 2023. https://arxiv.org/abs/2310.03368 Evaluating hallucinations in chinese large language models . Preprint, arXiv:2310.03368
arXiv 2023
-
[8]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . Preprint, arXiv:2110.14168
arXiv 2021
Show all 52 references
-
[9]
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. https://arxiv.org...
2024 arXiv
-
[10]
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. 2024. Does fine-tuning llms on new knowledge encourage hallucinations? arXiv preprint arXiv:2405.05904
2024 arXiv
-
[11]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[12]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[13]
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024. https://arxiv.org/abs/2402.14008 Olympiadbench: A challenging benchmark for promoting agi with olym...
2024 arXiv
-
[14]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. NeurIPS
2021
-
[15]
Jian Hu. 2025. Reinforce++: A simple and efficient approach for aligning large language models. arXiv preprint arXiv:2501.03262
2025 arXiv
-
[16]
Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung-Yeung Shum. 2025. https://arxiv.org/abs/2503.24290 Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model . Preprint, arXiv:2503.24290
2025 arXiv
-
[17]
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. 2024 a . Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In Proceedings of the IEEE/CVF ...
2024
-
[18]
Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu. 2024 b . O1 replication journey--part 2: Surpassing o1-preview through simple distillation, big progress or bitter lesson? arXiv preprint arXiv:2411.16489
2024 arXiv
-
[19]
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, and 1 others. 2022. Solving quantitative reasoning problems with language models. Advances in Neural Information Processin...
2022
-
[21]
Junyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024 b . The dawn after the dark: An empirical study on factuality hallucination in large language models. arXiv preprint arXiv:2401.03205
2024 arXiv
-
[22]
Shawn Li, Jiashu Qu, Yuxiao Zhou, Yuehan Qin, Tiankai Yang, and Yue Zhao. 2025 a . https://arxiv.org/abs/2503.06169 Treble counterfactual vlms: A causal approach to hallucination . Preprint, arXiv:2503.06169
2025 arXiv
-
[23]
Xuefeng Li, Haoyang Zou, and Pengfei Liu. 2025 b . Limr: Less is more for rl scaling. arXiv preprint arXiv:2502.11886
2025 arXiv
-
[24]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let's verify step by step. arXiv preprint arXiv:2305.20050
2023 arXiv
-
[25]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. https://doi.org/10.18653/v1/2022.acl-long.229 T ruthful QA : Measuring how models mimic human falsehoods . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...
2022 doi
-
[26]
Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. 2025. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://github.com/PraMamba/DeepScaleR. Notion Blog
2025
-
[27]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint
2022
-
[28]
OpenAI. 2024. https://cdn.openai.com/o1-system-card-20241205.pdf Openai o1 system card . Accessed: 2025-05-02
2024
-
[29]
OpenAI. 2025. https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf Openai o3 and o4-mini system card . Accessed: 2025-05-02
2025
-
[30]
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, and 1 others. 2023. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Comp...
2023
-
[31]
Qwen Team . 2025. QwQ-32B: Embracing the Power of Reinforcement Learning . https://qwenlm.github.io/blog/qwq-32b/. Accessed: 2025-05-02
2025
-
[32]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347
2017 arXiv
-
[33]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300
2024 arXiv
-
[34]
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. https://doi.org/10.1145/3689031.3696075 Hybridflow: A flexible and efficient rlhf framework . In Proceedings of the Twentieth European Conference on Compute...
2025
-
[35]
Taiwei Shi, Yiyang Wu, Linxin Song, Tianyi Zhou, and Jieyu Zhao. 2025. Efficient reinforcement finetuning via adaptive curriculum learning. arXiv preprint arXiv:2504.05520
2025 arXiv
-
[36]
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, and 1 others. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, pages 1--8
2025
-
[37]
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. 2024. Benchmarking hallucination in large language models based on unanswerable math word problem. In LREC/COLING
2024
-
[38]
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, and 1 others. 2023. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525
2023 arXiv
-
[39]
RUCAIBox STILL Team. 2025. https://github.com/RUCAIBox/Slow_Thinking_with_LLMs Still-3-1.5b-preview: Enhancing slow thinking abilities of small models through reinforcement learning
2025
-
[40]
M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das
S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. https://arxiv.org/abs/2401.01313 A comprehensive survey of hallucination mitigation techniques in large language models . Preprint, arXiv:2401.01313
2024 arXiv
-
[41]
Junyang Wang, Yuhang Wang, Guohai Xu, Jing Zhang, Yukai Gu, Haitao Jia, Jiaqi Wang, Haiyang Xu, Ming Yan, Ji Zhang, and 1 others. 2023. Amber: An llm-free multi-dimensional benchmark for mllms hallucination evaluation. arXiv preprint arXiv:2311.07397
2023 arXiv
-
[42]
O mer Faruk Akg \
Shangshang Wang, Julian Asilis, \"O mer Faruk Akg \"u l, Enes Burak Bilgin, Ollie Liu, and Willie Neiswanger. 2025 a . Tina: Tiny reasoning models via lora. arXiv preprint arXiv:2504.15777
2025 arXiv
-
[43]
Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Lucas Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, and 1 others. 2025 b . Reinforcement learning for reasoning in large language models with one training example. arXiv preprint arXiv:2504.20571
2025 arXiv
-
[44]
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le. 2023. Simple synthetic data reduces sycophancy in large language models. arXiv preprint arXiv:2308.03958
2023 arXiv
-
[45]
Xiaomi LLM-Core Team . 2025. https://github.com/XiaomiMiMo/MiMo Mimo: Unlocking the reasoning potential of language model – from pretraining to posttraining
2025
-
[46]
Wei Xiong, Jiarui Yao, Yuhui Xu, Bo Pang, Lei Wang, Doyen Sahoo, Junnan Li, Nan Jiang, Tong Zhang, Caiming Xiong, and Hanze Dong. 2025. https://arxiv.org/abs/2504.11343 A minimalist approach to llm reasoning: from rejection sampling to reinforce . Preprint, arXiv:2504.11343
2025 arXiv
-
[47]
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2024. Alignment for honesty. Advances in Neural Information Processing Systems, 37:63565--63598
2024
-
[48]
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don't know? arXiv preprint arXiv:2305.18153
2023 arXiv
-
[49]
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, and 1 others. 2025. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476
2025 arXiv
-
[50]
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, and 1 others. 2024. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proceedings of the IEEE/CVF Con...
2024
-
[51]
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. 2023. How language model hallucinations can snowball. arXiv preprint arXiv:2305.13534
2023 arXiv
-
[52]
Shaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz, Bryan Catanzaro, Andrew Tao, Qingyun Wu, Zhiding Yu, and Guilin Liu. 2025. https://arxiv.org/abs/2505.00024 Nemotron-research-tool-n1: Tool-using language models with reinforced reasoning . arXiv preprint arXiv:2505.00024
2025 arXiv
-
[53]
Andrew Zhao, Yiran Wu, Yang Yue, Tong Wu, Quentin Xu, Yang Yue, Matthieu Lin, Shenzhi Wang, Qingyun Wu, Zilong Zheng, and Gao Huang. 2025. https://arxiv.org/abs/2505.03335 Absolute zero: Reinforced self-play reasoning with zero data . Preprint, arXiv:2505.03335
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.