REVIEW 3 major objections 6 minor 8 cited by
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Long-CoT reasoning is no trustworthiness upgrade, survey finds.
desk verdict Useful map of reasoning trustworthiness literature, but the abstract's 'comparable or even greater' claim is undercut by the body's own benchmark-dependent evidence; revise, don't reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the survey's taxonomy of trustworthy reasoning: five dimensions (truthfulness, safety, robustness, fairness, privacy) crossed with two implementation paradigms (chain-of-thought prompting and end-to-end large reasoning models). Within that grid, the load-bearing object is the chain of thought itself—the generated intermediate reasoning steps that serve simultaneously as a faithfulness measurement target, a jailbreak surface, a backdoor trigger channel, a source of overthinking and underthinking, and a privacy side channel.
What would settle it
Run a standardized safety battery on the same reasoning models across many languages and attack templates and check whether open-source reasoning models remain consistently less safe than their base counterparts; if the rank order reverses across datasets, the pooled conclusion about comparable or greater vulnerabilities weakens.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that while chain-of-thought prompting and reasoning models can improve truthfulness through hallucination mitigation and support safety defenses such as guardrail models, the same reasoning capability produces new failure modes: reasoning models hallucinate in simple non-reasoning tasks, generate more harmful content after jailbreak, are less robust to small input perturbations, preserve or amplify bias, and leak more private information through their thinking traces. The synthesis, stated in the abstract, is that current reasoning models 'often suffer from comparable or even greater vulnerabilities' in safety, robustness, and privacy, with the
Load-bearing premise
The survey's cross-model conclusions assume that safety and robustness results measured on different benchmarks, languages, and attack templates can be combined into a single verdict; the paper itself reports cases where model rankings flip depending on the dataset.
Editorial extensions
If this is right
- Safety auditing of deployed reasoning models should monitor the thinking trace, not just the final answer, because several cited studies find the reasoning content is less safe than the output.
- Jailbreak defenses validated on chat models cannot be assumed to transfer to reasoning models; new attacks in this taxonomy specifically exploit reasoning traces, ciphers, or stepwise decomposition.
- Faithfulness evaluation needs standardized protocols; current intervention metrics can confound model strength with apparent unfaithfulness.
- Aligning reasoning models requires chain-of-thought-specific data and training, since safety reasoning must be taught rather than assumed to follow from general reasoning ability.
- Privacy and interpretability conflict: visible thinking traces improve transparency but also enable attribute inference and unlearning-recovery attacks.
Reading between the lines
- Not claimed by the paper, but the dataset-dependence of safety rankings suggests that single-number attack-success-rate scores can mislead; reporting per-category and per-language results would be a cheap test of which vulnerabilities are stable.
- Not claimed by the paper, but the finding that shortened reasoning improves harmlessness points to a concrete deployable intervention: force short reasoning at inference time and measure safety across languages.
- Not claimed by the paper, but if the safety tax scales with reasoning-length incentives, separating rewards for correctness from rewards for safety during reinforcement learning may decouple the trade-off.
- Not claimed by the paper, but if the multilingual vulnerabilities are a form of mismatched generalization, then training and evaluating safety in a few dominant languages systematically underestimates real-world deployment risk.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews the trustworthiness of chain-of-thought prompting and large reasoning models (LRMs) across five dimensions: truthfulness, safety, robustness, fairness, and privacy. It organizes recent work into a taxonomy, summarizes methods and findings in each area, and identifies open problems. The central claim, stated in the abstract and Introduction, is that reasoning techniques can improve some aspects of trustworthiness (e.g., hallucination mitigation, harmful-content detection, robustness) while cutting-edge reasoning models themselves often exhibit comparable or greater vulnerabilities in safety, robustness, and privacy than non-reasoning models. The paper positions itself as the first comprehensive survey of trustworthy reasoning and provides a public repository of related papers.
Significance. If the synthesis is reliable, the survey would be a useful resource for the AI-safety community. Its main strengths are the breadth of covered topics (five trustworthiness dimensions), the explicit separation of early CoT techniques from end-to-end reasoning models, and the identification of several concrete open problems, such as the need for standardized faithfulness metrics and more fine-grained benchmarks. The paper also highlights a genuinely important tension: improved reasoning capability does not automatically translate into improved trustworthiness. The structured taxonomy and the GitHub resource add practical value. However, the central comparative claim—that reasoning models 'often suffer from comparable or even greater vulnerabilities'—is only partially supported by the evidence the paper itself presents, because the underlying evaluations are heterogeneous and sometimes contradictory. The survey also contains several citation and attribution errors that undermine reliability, especially for a resource whose purpose is to guide readers to the correct literature.
major comments (3)
- [Abstract and §4.1] The abstract claims that reasoning models 'often suffer from comparable or even greater vulnerabilities' in safety, robustness, and privacy. This comparative claim is not stable under the evidence reported in §4.1. The paper's own third bullet explicitly states that 'pairwise safety ranks between models depend on datasets': AirBench finds DeepSeek-R1 safer than DeepSeek-V3, while CNSafe finds the opposite with an average ASR margin of 21.7%, and WildGuard Jailbreak vs CNSafe_RT again reverse the picture. Similar benchmark-dependence appears in robustness (§5.2: reasoning models beat non-reasoning models on CodeCrash but lose on Math-RoB and other perturbation benchmarks). Since these evaluations use different attack templates, languages, and model versions, pooling them into a single 'comparable or even greater' verdict is unsupported. The abstract and conclusion should either be qualifi
- [Introduction and §2.2] Several references are mischaracterized or misassigned. In the Introduction, reference [12] is described as 'a related survey [12] provided valuable discussions on safety-related aspects,' but [12] is 'Safety Reasoning with Guidelines' (ICML 2025), an original research paper, not a survey. In §2.2, the sentence 'zero-shot-CoT [16]' attributes zero-shot CoT to Wei et al. [16]; the correct reference for zero-shot CoT is Kojima et al. [17]. In addition, 'few-shot-CoT [19]' in the same paragraph is wrong: [19] is the Chain-of-Scrutiny backdoor-detection paper, not the few-shot CoT paper. For a survey whose purpose is to organize and signpost the literature, these citation errors are material and should be corrected throughout.
- [§3.2.2 and §8] The faithfulness section reports directly contradictory conclusions (e.g., 'larger models are generally more faithful' [85,78] vs 'models with higher accuracy tend to exhibit lower faithfulness' [79,86]; 'length penalties may result in unfaithful responses' [81] vs 'unfaithful CoTs are usually longer' [82]). The survey notes these contradictions and calls for standardized metrics, which is appropriate. However, the synthesis does not identify which differences are likely due to evaluation methodology (e.g., intervention type, task difficulty, model family) versus genuine model properties. Since the paper explicitly lists 'standard measurements of faithfulness' as a future direction, it should at least organize the existing evidence around the methodological axes that are already discussed in §3.2.1. Without that, the 'factors that influence faithfulness' subsection is more a list of conf
minor comments (6)
- [§2.2] The citation typo 'zero-shot-CoT [16]' should read [17]; the nearby 'few-shot-CoT [19]' should read [16] (or the intended reference). Please check all citation numbers against the bibliography.
- [§3.2.1] Typo: 'Lakage-Adjusted Simulatability' should be 'Leakage-Adjusted Simulatability'; 'Paulet al.' should be 'Paul et al.'
- [§4.3.3] Typo: 'convolution neural network' should be 'convolutional neural network.'
- [Figure 2] The taxonomy figure is dense and the font is small. Consider grouping by sub-theme or providing an accompanying table to improve readability.
- [References] Reference formats are inconsistent: some entries include only arXiv identifiers and no publication venue, while others include venue names. For a survey, adding DOIs or stable URLs would improve usability.
- [§6] The fairness section is comparatively short and does not discuss the interaction of fairness with reasoning-model training (e.g., RLVR) beyond citing a few benchmarks. A brief discussion of open problems in this specific intersection would strengthen the survey.
Circularity Check
No material circularity: survey synthesis rests on external literature; self-citations are minor and not load-bearing.
full rationale
This is a survey paper rather than a derivation; its central claim—that reasoning techniques can improve some aspects of trustworthiness while reasoning models themselves show comparable or greater vulnerabilities—is a synthesis of many independent empirical studies, not the output of a fitted model or a self-referential equation. No step was found where a parameter is fit to a subset of data and then renamed a prediction, and no uniqueness theorem from the authors' own prior work is invoked to force a conclusion. The authors do cite several of their own papers ([260], [304], [305], [340], [344], [345]), but in all cases these are background or supporting citations within broader enumerations of related work; the central conclusions are supported by numerous third-party benchmarks and studies. The paper even explicitly acknowledges benchmark-dependent disagreements in Section 4.1 ('Pairwise safety ranks between models depend on datasets'), which is a limitation on comparability rather than a circularity. Under the hard rules requiring a specific quoted reduction for a circularity finding, no such reduction can be exhibited. The score of 2 reflects the presence of minor self-citations that are not load-bearing, not any actual circular reasoning.
Assumptions & free parameters
assumptions (3)
- domain assumption The five categories (truthfulness, safety, robustness, fairness, privacy) constitute an adequate decomposition of trustworthiness.
- domain assumption The summarized findings in the cited papers are reliable and representative of the literature up to June 2025.
- domain assumption Conflicting evaluation results across benchmarks can be reconciled into a cross-model conclusion.
Cite this review
Pith. "Pith review of A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models." pith.science (2026). https://pith.science/paper/P5SZ5446
@misc{pith2026250903871,
author = {Pith},
title = {Pith review of: A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/P5SZ5446}},
note = {Machine review of arXiv:2509.03871}
}
read the original abstract
The development of Long-CoT reasoning has advanced LLM performance across various tasks, including language understanding, complex problem solving, and code generation. This paradigm enables models to generate intermediate reasoning steps, thereby improving both accuracy and interpretability. However, despite these advancements, a comprehensive understanding of how CoT-based reasoning affects the trustworthiness of language models remains underdeveloped. In this paper, we survey recent work on reasoning models and CoT techniques, focusing on five core dimensions of trustworthy reasoning: truthfulness, safety, robustness, fairness, and privacy. For each aspect, we provide a clear and structured overview of recent studies in chronological order, along with detailed analyses of their methodologies, findings, and limitations. Future research directions are also appended at the end for reference and discussion. Overall, while reasoning techniques hold promise for enhancing model trustworthiness through hallucination mitigation, harmful content detection, and robustness improvement, cutting-edge reasoning models themselves often suffer from comparable or even greater vulnerabilities in safety, robustness, and privacy. By synthesizing these insights, we hope this work serves as a valuable and timely resource for the AI safety community to stay informed on the latest progress in reasoning trustworthiness. A full list of related papers can be found at \href{https://github.com/ybwang119/Awesome-reasoning-safety}{https://github.com/ybwang119/Awesome-reasoning-safety}.
Figures
Forward citations
Cited by 8 Pith papers
-
Risky Business: Measuring The Faithfulness-Safety Tension
Faithful reasoning and safety pull in opposite directions in current reasoning models, and the two behaviors are controlled by anti-correlated internal vectors that can be steered independently.
-
Where Do CoT Training Gains Land in LLM based Agents?
CoT training in LLM agents improves prompt-action quality more than the advantage of generated reasoning, and selectively masking action supervision improves out-of-domain generalization.
-
Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries
Swapping the reasoning trace prefill on unlearned weights can replicate or reverse the parser-split bypass gap, showing that the gap alone does not identify or rule out weight-level memorization.
-
Pause or Fabricate? Training Language Models for Grounded Reasoning
GRIL uses stage-specific RL rewards to train LLMs to detect missing premises, pause proactively, and resume grounded reasoning after clarification, yielding up to 45% better premise detection and 30% higher task succe...
-
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
PreRL applies reward-driven updates to P(y) in pre-train space, uses Negative Sample Reinforcement to prune bad reasoning paths and boost reflection, and combines with standard RL in Dual Space RL to outperform baseli...
-
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
TRACE-RPS drops LLM attribute inference accuracy from around 50% to below 5% via fine-grained anonymization plus a two-stage rejection optimization.
-
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
Faithful Warm-Start pre-training on causally consistent vision-language samples improves accuracy, stabilizes RL, and reduces unsupported reasoning in VLMs.
-
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
A 16-factor structured prompt framework strengthens CoT reasoning in LLMs for security analysis, yielding up to 40% reasoning gains in smaller models and stable accuracy improvements validated by human raters with Coh...
Reference graph
Works this paper leans on
-
[12]
Safety Reasoning with Guidelines
Haoyu Wang, Zeyu Qin, Li Shen, Xueqian Wang, Dacheng Tao, and Minhao Cheng. Safety Reasoning with Guidelines. In Proc. ICML, 2025
2025
-
[16]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. In Proc. NeurIPS, 2022
2022
-
[17]
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Proc. NeurIPS, 2022
2022
-
[19]
Chain-of-scrutiny: Detecting backdoor attacks for large language models
Xi Li, Yusen Zhang, Renze Lou, Chen Wu, and Jiaqi Wang. Chain-of-scrutiny: Detecting backdoor attacks for large language models. arXiv preprint arXiv:2406.05948, 2024
arXiv 2024
-
[81]
Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025
James Chua and Owain Evans. Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025
2025
-
[82]
Reasoning Models Don’t Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner Fabien Roger Vlad Mikulik, Sam Bowman, Jan Leike Jared Kaplan, et al. Reasoning Models Don’t Always Say What They Think. Anthropic Research, 2025
2025
-
[1]
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. Safechain: Safety of language models with long chain-of-thought reasoning capabilities. arXiv preprint arXiv:2502.12025, 2025
arXiv 2025
-
[2]
Chengda Lu, Xiaoyu Fan, Yu Huang, Rongwu Xu, Jijie Li, and Wei Xu. Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking? arXiv preprint arXiv:2505.17650, 2025
arXiv 2025
Show all 299 references
-
[3]
Towards understanding the safety boundaries of deepseek models: Evaluation and findings
Zonghao Ying, Guangyi Zheng, Yongxin Huang, Deyue Zhang, Wenxin Zhang, Quanchen Zou, Aishan Liu, Xianglong Liu, and Dacheng Tao. Towards understanding the safety boundaries of deepseek models: Evaluation and findings. arXiv preprint arXiv:2503.15092, 2025
2025 arXiv
-
[4]
A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment
Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, et al. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585, 2025
2025 arXiv
-
[5]
Attacks, defenses and evaluations for llm conversation safety: A survey
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. Attacks, defenses and evaluations for llm conversation safety: A survey. In Proc. NAACL, 2024
2024
-
[6]
Large language model safety: A holistic survey
Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li, Yongqi Leng, Renren Jin, Chuang Liu, Xinwei Wu, Zishan Guo, Linhao Yu, et al. Large language model safety: A holistic survey. arXiv preprint arXiv:2412.17686, 2024
2024 arXiv
-
[7]
Towards reasoning era: A survey of long chain-of-thought for reasoning large language models
Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wangxiang Che. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models. arXiv preprint arXiv:2503.09567, 2025
2025 arXiv
-
[8]
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv preprint arXiv:2501.09686, 2025
2025 arXiv
-
[9]
A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, et al. A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond. arXiv preprint arXiv:2503.21614, 2025
2025
-
[10]
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, et al. Stop overthinking: A survey on efficient reasoning for large language models. arXiv preprint arXiv:2503.16419, 2025
2025 arXiv
-
[11]
Efficient reasoning models: A survey
Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. Efficient reasoning models: A survey. arXiv preprint arXiv:2504.10903, 2025
2025
-
[13]
Gpt-4 technical report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[14]
The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Anthropic. The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
2024
-
[15]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[18]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In Proc. NeurIPS, 2020
2020
-
[20]
Openai o1 system card
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024
2024 arXiv
-
[21]
Openr: An open source framework for advanced reasoning with large language models
Jun Wang, Meng Fang, Ziyu Wan, Muning Wen, Jiachen Zhu, Anjie Liu, Ziqin Gong, Yan Song, Lei Chen, Lionel M Ni, et al. Openr: An open source framework for advanced reasoning with large language models. arXiv preprint arXiv:2410.09671, 2024
2024 arXiv
-
[22]
O1 Replication Journey: A Strategic Progress Report–Part 1
Yiwei Qin, Xuefeng Li, Haoyang Zou, Yixiu Liu, Shijie Xia, Zhen Huang, Yixin Ye, Weizhe Yuan, Hector Liu, Yuanzhi Li, et al. O1 Replication Journey: A Strategic Progress Report–Part 1. arXiv preprint arXiv:2410.18982, 2024
2024 arXiv
-
[23]
O1 Replication Journey–Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? arXiv preprint arXiv:2411.16489, 2024
Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu. O1 Replication Journey–Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? arXiv preprint arXiv:2411.16489, 20...
2024 arXiv
-
[24]
O1 Replication Journey–Part 3: Inference-time Scaling for Medical Reasoning
Zhongzhen Huang, Gui Geng, Shengyi Hua, Zhen Huang, Haoyang Zou, Shaoting Zhang, Pengfei Liu, and Xiaofan Zhang. O1 Replication Journey–Part 3: Inference-time Scaling for Medical Reasoning. arXiv preprint arXiv:2501.06458, 2025
2025 arXiv
-
[25]
Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning
Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, et al. Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning. arXiv preprint arXiv:2410.02884, 2024
2024 arXiv
-
[26]
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in...
2012
-
[27]
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[28]
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. In Proc. NeurIPS D&B Track, 2021
2021
-
[29]
MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline
Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline. In Findings of Proc. ACL, page 905–924, 2024
2024
-
[30]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Proc. NeurIPS, 2023
2023
-
[31]
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[32]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[33]
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv preprint arXiv:2502.05171, 2025
2025 arXiv
-
[34]
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space. In ICLR Workshop on LLM Reason and Plan, 2024
2024
-
[35]
Deepseek-v3 technical report
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[36]
Qwen2.5 technical report
A Yang Qwen, Baosong Yang, B Zhang, B Hui, B Zheng, B Yu, Chengpeng Li, D Liu, F Huang, H Wei, et al. Qwen2.5 technical report. arXiv preprint, 2024
2024
-
[37]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv e-prints, pages arXiv–2407, 2024
2024
-
[38]
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022
2022 arXiv
-
[39]
Star: Self-taught reasoner bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D Goodman. Star: Self-taught reasoner bootstrapping reasoning with reasoning. In Proc. NeurIPS, volume 1126, 2024
2024
-
[40]
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In Proc. ICLR, 2023
2023
-
[41]
T \" ulu 3: Pushing frontiers in open language model post-training
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al. T \" ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024
2024 arXiv
-
[42]
Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation
Youssef Mroueh. Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation. arXiv preprint arXiv:2503.06639, 2025
2025
-
[43]
Perception, reason, think, and plan: A survey on large multimodal reasoning models
Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. Perception, reason, think, and plan: A survey on large multimodal reasoning models. arXiv preprint arXiv:2505.04921, 2025
2025 arXiv
-
[44]
Multimodal chain-of- thought reasoning: A comprehensive survey
Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. Multimodal chain-of- thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605, 2025
2025 arXiv
-
[45]
Multimodal chain-of-thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research, 2024
2024
-
[46]
Video-of-thought: Step-by-step video reasoning from perception to cognition
Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Proc. ICML, 2024
2024
-
[47]
Thinking before looking: Improving multimodal llm reasoning via mitigating visual hallucination
Haojie Zheng, Tianyang Xu, Hanchi Sun, Shu Pu, Ruoxi Chen, and Lichao Sun. Thinking before looking: Improving multimodal llm reasoning via mitigating visual hallucination. arXiv preprint arXiv:2411.12591, 2024. 25 References A Preprint
2024 arXiv
-
[48]
Llava-o1: Let vision language models reason step-by-step
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan. Llava-o1: Let vision language models reason step-by-step. arXiv preprint arXiv:2411.10440, 2024
2024 arXiv
-
[49]
Llamav-o1: Rethinking step-by-step visual reasoning in llms
Omkar Thawakar, Dinura Dissanayake, Ketan More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, et al. Llamav-o1: Rethinking step-by-step visual reasoning in llms. arXiv preprint arXiv:2501.06186, 2025
2025 arXiv
-
[50]
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? arXiv preprint arXiv:2501.11284, 2025
Haotian Xu, Xing Wu, Weinong Wang, Zhongzhi Li, Da Zheng, Boyuan Chen, Yi Hu, Shijia Kang, Jiaming Ji, Yingying Zhang, et al. RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? arXiv preprint arXiv:2501.11284, 2025
2025 arXiv
-
[51]
Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search
Huanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang, Yibo Wang, Shunyu Liu, Yingjie Wang, Yuxin Song, Haocheng Feng, Li Shen, et al. Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search. arXiv preprint arXiv:2412.18319, 2024
2024 arXiv
-
[52]
Improve vision language model chain-of-thought reasoning
Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, and Yiming Yang. Improve vision language model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198, 2024
2024 arXiv
-
[53]
Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. In Proc. CVPR, pages 9062–9072, 2025
2025
-
[54]
Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale
Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xiang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237, 2024
2024 arXiv
-
[55]
Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification
Linzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, and Wentao Zhang. Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification. arXiv preprint arXiv:2502.13383, 2025
2025 arXiv
-
[56]
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking. arXiv preprint arXiv:2501.01306, 2025
2025 arXiv
-
[57]
HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning
Shayan Ali Akbar, Md Mosharaf Hossain, Tess Wood, Si-Chi Chin, Erica M Salinas, Victor Alvarez, and Erwin Cornejo. HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning. In Proc. EMNLP, pages 15020– 15037, 2024
2024
-
[58]
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
Ron Eliav, Arie Cattan, Eran Hirsch, Shahaf Bassan, Elias Stengel-Eskin, Mohit Bansal, and Ido Dagan. CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection. arXiv preprint arXiv:2506.05243, 2025
2025 arXiv
-
[59]
Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language- Models
Zikai Xie. Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language- Models. arXiv preprint arXiv:2408.05093, 2024
2024 arXiv
-
[60]
Grounded chain-of- thought for multimodal large language models
Qiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang, Baiyang Song, Xiaoshuai Sun, and Rongrong Ji. Grounded chain-of- thought for multimodal large language models. arXiv preprint arXiv:2503.12799, 2025
2025 arXiv
-
[61]
CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation
Yue Jiang, Jiawei Chen, Dingkang Yang, Mingcheng Li, Shunli Wang, Tong Wu, Ke Li, and Lihua Zhang. CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation. In Proc. ICASSP, 2025
2025
-
[62]
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Bowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, and Lei Zhang. MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM. arXiv preprint arXiv:2505.24238, 2025
2025 arXiv
-
[63]
The Hallucination Tax of Reinforcement Finetuning.arXiv preprint arXiv:2505.13988, 2025
Linxin Song, Taiwei Shi, and Jieyu Zhao. The Hallucination Tax of Reinforcement Finetuning.arXiv preprint arXiv:2505.13988, 2025
2025 arXiv
-
[64]
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou, Xin Eric Wang, Yuyin Zhou, and Sheng Liu. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models. arXiv preprint arXiv:2505.21523, 2025
2025 arXiv
-
[65]
Are Reasoning Models More Prone to Hallucination? arXiv preprint arXiv:2505.23646, 2025
Zijun Yao, Yantao Liu, Yanxu Chen, Jianhui Chen, Junfeng Fang, Lei Hou, Juanzi Li, and Tat-Seng Chua. Are Reasoning Models More Prone to Hallucination? arXiv preprint arXiv:2505.23646, 2025
2025 arXiv
-
[66]
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, and Samuel J Bell. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions. arXiv preprint arXiv:2506.09038, 2025
2025 arXiv
-
[67]
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
Haolang Lu, Yilian Liu, Jingxin Xu, Guoshun Nan, Yuanlong Yu, Zhican Chen, and Kun Wang. Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models. arXiv preprint arXiv:2505.13143, 2025
2025
-
[68]
The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models
Junyi Li and Hwee Tou Ng. The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models. arXiv preprint arXiv:2505.24630, 2025
2025
-
[69]
Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning
Dang Hoang Anh, Vu Tran, and Le Minh Nguyen. Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning. In JSAI International Symposium on Artificial Intelligence, pages 179–195. Springer, 2025
2025
-
[70]
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
Zhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang, and Jun Xu. Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective. arXiv preprint arXiv:2505.12886, 2025
2025 arXiv
-
[71]
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
Dadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He, Haoran Li, Yumeng Wang, et al. Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models. arXiv preprint arXiv:2506.17114, 2025
2025
-
[72]
Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning
Ruosen Li, Ziming Luo, and Xinya Du. Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning. arXiv preprint arXiv:2410.06304, 2024. 26 References A Preprint
2024
-
[73]
Reasoning Models Know When They’re Right: Probing Hidden States for Self-Verification
Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, and He He. Reasoning Models Know When They’re Right: Probing Hidden States for Self-Verification. arXiv preprint arXiv:2504.05419, 2025
2025 arXiv
-
[74]
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
Changyue Wang, Weihang Su, Qingyao Ai, and Yiqun Liu. Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models. arXiv preprint arXiv:2506.04832, 2025
2025
-
[75]
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702, 2023
2023 arXiv
-
[76]
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Proc. NeurIPS, 2023
2023
-
[77]
Measuring faithfulness of chains of thought by unlearning reasoning steps
Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasović, and Yonatan Belinkov. Measuring faithfulness of chains of thought by unlearning reasoning steps. arXiv preprint arXiv:2502.14829, 2025
2025
-
[78]
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Zidi Xiong, Chen Shan, Zhenting Qi, and Himabindu Lakkaraju. Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models. arXiv preprint arXiv:2505.13774, 2025
2025 arXiv
-
[79]
Chain-of-Thought Unfaithfulness as Disguised Accuracy
Oliver Bentham, Nathan Stringham, and Ana Marasovic. Chain-of-Thought Unfaithfulness as Disguised Accuracy. Transac- tions on Machine Learning Research, 2024
2024
-
[80]
Chain-of- thought reasoning in the wild is not always faithful
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy. Chain-of- thought reasoning in the wild is not always faithful. arXiv preprint arXiv:2503.08679, 2025
2025 arXiv
-
[83]
Towards faithful chain-of-thought: Large language models are bridging reasoners
Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. Towards faithful chain-of-thought: Large language models are bridging reasoners. arXiv preprint arXiv:2405.18915, 2024
2024 arXiv
-
[84]
Faithfulness vs
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. Faithfulness vs. plausibility: On the (un) reliability of explanations from large language models. arXiv preprint arXiv:2402.04614, 2024
2024 arXiv
-
[85]
How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc
Guangsheng Bao, Hongbo Zhang, Cunxiang Wang, Linyi Yang, and Yue Zhang. How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc. COLING, 2024
2024
-
[86]
On the difficulty of faithful chain-of-thought reasoning in large language models
Sree Harsha Tanneru, Dan Ley, Chirag Agarwal, and Himabindu Lakkaraju. On the difficulty of faithful chain-of-thought reasoning in large language models. In ICML Workshop on TiFA, 2024
2024
-
[87]
On the impact of fine-tuning on chain-of-thought reasoning
Elita Lobo, Chirag Agarwal, and Himabindu Lakkaraju. On the impact of fine-tuning on chain-of-thought reasoning. In Proc. NAACL, 2025
2025
-
[88]
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. In Findings of Proc. EMNLP, pages 15012–15032, 2024
2024
-
[89]
Faithful logical reasoning via symbolic chain-of-thought
Jundong Xu, Hao Fei, Liangming Pan, Qian Liu, Mong-Li Lee, and Wynne Hsu. Faithful logical reasoning via symbolic chain-of-thought. In Proc. ACL, 2024
2024
-
[90]
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamil˙e Lukoši¯ut˙e, et al. Question decomposition improves the faithfulness of model-generated reasoning. arXiv preprint arXiv:2307.11768, 2023
2023 arXiv
-
[91]
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. Faithful chain-of-thought reasoning. In Proc. IJCNLP-AACL, 2023
2023
-
[92]
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning. In Proc. EMNLP, 2023
2023
-
[93]
FLARE: Faithful Logic-Aided Reasoning and Exploration
Erik Arakelyan, Pasquale Minervini, Pat Verga, Patrick Lewis, and Isabelle Augenstein. FLARE: Faithful Logic-Aided Reasoning and Exploration. arXiv preprint arXiv:2410.11900, 2024
2024
-
[94]
CoMAT: Chain of mathematically annotated thought improves mathematical reasoning
Joshua Ong Jun Leang, Aryo Pradipta Gema, and Shay B Cohen. CoMAT: Chain of mathematically annotated thought improves mathematical reasoning. arXiv preprint arXiv:2410.10336, 2024
2024
-
[95]
Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering
Jiawei Wang, Da Cao, Shaofei Lu, Zhanchang Ma, Junbin Xiao, and Tat-Seng Chua. Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering. In Proc. MM, pages 4331–4340, 2024
2024
-
[96]
Fact: Teaching mllms with faithful, concise and transferable rationales
Minghe Gao, Shuang Chen, Liang Pang, Yuan Yao, Jisheng Dang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang, and Tat-Seng Chua. Fact: Teaching mllms with faithful, concise and transferable rationales. In Proc. MM, pages 846–855, 2024
2024
-
[97]
Markovian Transformers for Informative Language Modeling
Scott Viteri, Max Lamparth, Peter Chatain, and Clark Barrett. Markovian Transformers for Informative Language Modeling. arXiv preprint arXiv:2404.18988, 2024. 27 References A Preprint
2024
-
[98]
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Limin Han, Jiaojiao Zhao, Beibei Huang, Zhenhong Long, Junting Guo, Meijuan An, Rongjia Du, et al. Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts. arXiv preprint arXiv:2503.16529, 2025
2025 arXiv
-
[99]
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
Miguel Romero-Arjona, Pablo Valle, Juan C Alonso, Ana B Sánchez, Miriam Ugarte, Antonia Cazalilla, Vicente Cambrón, José A Parejo, Aitor Arrieta, and Sergio Segura. Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives. arXiv preprint arXiv:2503.10192, 2025
2025 arXiv
-
[100]
The hidden risks of large reasoning models: A safety assessment of r1
Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Shreedhar Jangam, Jayanth Srinivasa, Gaowen Liu, Dawn Song, and Xin Eric Wang. The hidden risks of large reasoning models: A safety assessment of r1. arXiv preprint arXiv:2502.12659, 2025
2025
-
[101]
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
Ang Li, Yichuan Mo, Mingjie Li, Yifei Wang, and Yisen Wang. Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning. arXiv preprint arXiv:2502.09673, 2025
2025 arXiv
-
[102]
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, and Kaiyu Huang. Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model. arXiv preprint arXiv:2505.06538, 2025
2025
-
[103]
Evaluating Security Risk in DeepSeek and Other Frontier Reasoning Models
Paul Kassianik and Amin Karbasi. Evaluating Security Risk in DeepSeek and Other Frontier Reasoning Models. Cisco, https://blogs. cisco. com/security/evaluating-security-risk-in-deepseek-and-other-frontier-reasoningmodels, 2025
2025
-
[104]
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
Arjun Krishna, Aaditya Rastogi, and Erick Galinkin. Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models. arXiv preprint arXiv:2506.13726, 2025
2025 arXiv
-
[105]
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
Christina Q Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Team, and Julian Michael. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety. arXiv preprint arXiv:2506.14922, 2025
2025 arXiv
-
[106]
Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
Yihe Fan, Wenqi Zhang, Xudong Pan, and Min Yang. Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems. arXiv preprint arXiv:2505.17815, 2025
2025
-
[107]
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Wenbo Su, Xiaoyong Zhu, et al. Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models. arXiv preprint arXiv:2505.19690, 2025
2025 arXiv
-
[108]
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
Xiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou, Weichen Zhang, Dongrui Liu, Lu Sheng, and Jing Shao. IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks. arXiv preprint arXiv:2506.16402, 2025
2025
-
[109]
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
Junfeng Fang, Yukai Wang, Ruipeng Wang, Zijun Yao, Kun Wang, An Zhang, Xiang Wang, and Tat-Seng Chua. SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models. arXiv preprint arXiv:2504.08813, 2025
2025 arXiv
-
[110]
Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over foundational capabilities
Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che, Tat-Seng Chua, and Ting Liu. Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over foundational capabilities. arXiv preprint arXiv:...
2025
-
[111]
DeepSeek-R1 Thoughtology: Let’s think about LLM Reasoning
Sara Vera Marjanović, Arkil Patel, Vaibhav Adlakha, Milad Aghajohari, Parishad BehnamGhader, Mehar Bhatia, Aditi Khandelwal, Austin Kraft, Benno Krojer, Xing Han Lù, et al. DeepSeek-R1 Thoughtology: Let’s think about LLM Reasoning. arXiv preprint arXiv:2504.07128, 2025
2025
-
[112]
Adversarial Reasoning at Jailbreaking Time
Mahdi Sabbaghi, Paul Kassianik, George Pappas, Yaron Singer, Amin Karbasi, and Hamed Hassani. Adversarial Reasoning at Jailbreaking Time. In Proc. ICML, 2025
2025
-
[113]
Enhancing Adversarial Attacks through Chain of Thought
Jingbo Su. Enhancing Adversarial Attacks through Chain of Thought. arXiv preprint arXiv:2410.21791, 2024
2024 arXiv
-
[114]
Reasoning-augmented conversation for multi-turn jailbreak attacks on large language models
Zonghao Ying, Deyue Zhang, Zonglei Jing, Yisong Xiao, Quanchen Zou, Aishan Liu, Siyuan Liang, Xiangzheng Zhang, Xianglong Liu, and Dacheng Tao. Reasoning-augmented conversation for multi-turn jailbreak attacks on large language models. arXiv preprint arXiv:2502.11054, 2025
2025 arXiv
-
[115]
Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models
Wenhan Chang, Tianqing Zhu, Yu Zhao, Shuangyong Song, Ping Xiong, Wanlei Zhou, and Yongxiang Li. Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models. arXiv preprint arXiv:2505.17519, 2025
2025
-
[116]
competency
Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, and Chitta Baral. When “competency" in reasoning opens the door to vulnerability: Jailbreaking llms via novel complex ciphers. arXiv preprint arXiv:2402.10601, 2024
2024
-
[117]
H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking
Martin Kuo, Jianyi Zhang, Aolin Ding, Qinsi Wang, Louis DiValentin, Yujia Bao, Wei Wei, Hai Li, and Yiran Chen. H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash think...
2025 arXiv
-
[118]
A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos
Yang Yao, Xuan Tong, Ruofan Wang, Yixu Wang, Lujundong Li, Liang Liu, Yan Teng, and Yingchun Wang. A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos. arXiv preprint arXiv:2502.15806, 2025
2025 arXiv
-
[119]
AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Models
Jiacheng Liang, Tanqiu Jiang, Yuhui Wang, Rongyi Zhu, Fenglong Ma, and Ting Wang. AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Models. arXiv preprint arXiv:2505.10846, 2025
2025 arXiv
-
[120]
Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers
Viet-Anh Nguyen, Shiqian Zhao, Gia Dao, Runyi Hu, Yi Xie, and Luu Anh Tuan. Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers. arXiv preprint arXiv:2505.16241, 2025
2025 arXiv
-
[121]
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, and Lap-Pui Chau. Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models. arXiv preprint arXiv:2504.05050, 2025
2025 arXiv
-
[122]
RRTL: Red Teaming Reasoning Large Language Models in Tool Learning.arXiv preprint arXiv:2505.17106, 2025
Yifei Liu, Yu Cui, and Haibin Zhang. RRTL: Red Teaming Reasoning Large Language Models in Tool Learning.arXiv preprint arXiv:2505.17106, 2025. 28 References A Preprint
2025 arXiv
-
[123]
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
Bingrui Sima, Linhua Cong, Wenxuan Wang, and Kun He. VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models. arXiv preprint arXiv:2505.19684, 2025
2025 arXiv
-
[124]
HauntAttack: When Attack Follows Reasoning as a Shadow
Jingyuan Ma, Rui Li, Zheng Li, Junfeng Liu, Lei Sha, and Zhifang Sui. HauntAttack: When Attack Follows Reasoning as a Shadow. arXiv preprint arXiv:2506.07031, 2025
2025 arXiv
-
[125]
GuardReasoner: Towards Reasoning-based LLM Safeguards
Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, and Bryan Hooi. GuardReasoner: Towards Reasoning-based LLM Safeguards. arXiv preprint arXiv:2501.18492, 2025
2025
-
[126]
X-Guard: Multilingual guard agent for content moderation
Bibek Upadhayay, Vahid Behzadan, et al. X-Guard: Multilingual guard agent for content moderation. arXiv preprint arXiv:2504.08848, 2025
2025 arXiv
-
[127]
Yahan Yang, Soham Dan, Shuo Li, Dan Roth, and Insup Lee. MR. Guard: Multilingual Reasoning Guardrail using Curriculum Learning. arXiv preprint arXiv:2504.15241, 2025
2025
-
[128]
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards.arXiv preprint arXiv:2506.07736, 2025
Jingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang, and Tat-Seng Chua. RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards.arXiv preprint arXiv:2506.07736, 2025
2025
-
[129]
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
Makesh Narsimhan Sreedhar, Traian Rebedea, and Christopher Parisien. Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models. arXiv preprint arXiv:2505.20087, 2025
2025 arXiv
-
[130]
R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
Mintong Kang and Bo Li. R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning. In Proc. ICLR, 2025
2025
-
[132]
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
Shiyao Cui, Qinglin Zhang, Xuan Ouyang, Renmiao Chen, Zhexin Zhang, Yida Lu, Hongning Wang, Han Qiu, and Minlie Huang. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs. arXiv preprint arXiv:2505.14035, 2025
2025 arXiv
-
[133]
Guardreasoner-vl: Safeguarding vlms via reinforced reasoning
Yue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang, Xinfeng Li, Kun Wang, Junfeng Fang, et al. Guardreasoner-vl: Safeguarding vlms via reinforced reasoning. arXiv preprint arXiv:2505.11049, 2025
2025 arXiv
-
[134]
GuardAgent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning
Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, et al. GuardAgent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning. arXiv preprint arXiv:2406.09187, 2024
2024 arXiv
-
[135]
ShieldAgent: Shielding agents via verifiable safety policy reasoning
Zhaorun Chen, Mintong Kang, and Bo Li. ShieldAgent: Shielding agents via verifiable safety policy reasoning. arXiv preprint arXiv:2503.22738, 2025
2025
-
[136]
Unified multimodal chain-of- thought reward model through reinforcement fine-tuning
Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, and Jiaqi Wang. Unified multimodal chain-of- thought reward model through reinforcement fine-tuning. arXiv preprint arXiv:2505.03318, 2025
2025
-
[137]
Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning
Fengjun Pan, Anh Tuan Luu, and Xiaobao Wu. Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning. arXiv preprint arXiv:2506.08477, 2025
2025
-
[138]
Effectively Controlling Reasoning Models through Thinking Intervention
Tong Wu, Chong Xiang, Jiachen T Wang, and Prateek Mittal. Effectively Controlling Reasoning Models through Thinking Intervention. arXiv preprint arXiv:2503.24370, 2025
2025 arXiv
-
[139]
Adversarial Manipulation of Reasoning Models using Internal Representations
Kureha Yamaguchi, Benjamin Etheridge, and Andy Arditi. Adversarial Manipulation of Reasoning Models using Internal Representations. In ICML 2025 Workshop on Reliable and Responsible Foundation Models, 2025
2025
-
[140]
Trading inference-time compute for adversarial robustness
Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak, Stephanie Lin, Sam Toyer, Yaodong Yu, Rachel Dias, Eric Wallace, Kai Xiao, Johannes Heidecke, et al. Trading inference-time compute for adversarial robustness. arXiv preprint arXiv:2501.18841, 2025
2025 arXiv
-
[141]
Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance
Ruizhong Qiu, Gaotang Li, Tianxin Wei, Jingrui He, and Hanghang Tong. Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance. arXiv preprint arXiv:2506.06444, 2025
2025 arXiv
-
[142]
Mixture of insightful experts (mote): The synergy of thought chains and expert mixtures in self-alignment
Zhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, et al. Mixture of insightful experts (mote): The synergy of thought chains and expert mixtures in self-alignment. arXiv preprint arXiv:2405.00557, 2024
2024 arXiv
-
[143]
Backtracking improves generation safety
Yiming Zhang, Jianfeng Chi, Hailey Nguyen, Kartikeya Upasani, Daniel M Bikel, Jason Weston, and Eric Michael Smith. Backtracking improves generation safety. In Proc. ICLR, 2025
2025
-
[144]
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
Xianglin Yang, Gelei Deng, Jieming Shi, Tianwei Zhang, and Jin Song Dong. Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning. arXiv preprint arXiv:2501.19180, 2025
2025
-
[145]
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
Junda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin, and Lei Sha. Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking. arXiv preprint arXiv:2502.12970, 2025
2025
-
[146]
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
Yuyou Zhang, Miao Li, William Han, Yihang Yao, Zhepeng Cen, and Ding Zhao. Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety. arXiv preprint arXiv:2503.05021, 2025
2025
-
[147]
ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimization
Kehua Feng, Keyan Ding, Jing Yu, Menghan Li, Yuhao Wang, Tong Xu, Xinda Wang, Qiang Zhang, and Huajun Chen. ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimization. arXiv preprint arXiv:2504.02725, 2025. 29 References A Preprint
2025
-
[148]
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
Yutao Mou, Yuxiao Luo, Shikun Zhang, and Wei Ye. SaRO: Enhancing LLM Safety through Reasoning-based Alignment. arXiv preprint arXiv:2504.09420, 2025
2025 arXiv
-
[149]
Reasoning as an Adaptive Defense for Safety.arXiv preprint arXiv:2507.00971, 2025
Taeyoun Kim, Fahim Tajwar, Aditi Raghunathan, and Aviral Kumar. Reasoning as an Adaptive Defense for Safety.arXiv preprint arXiv:2507.00971, 2025
2025
-
[150]
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Changyue Jiang, Xudong Pan, and Min Yang. Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction. arXiv preprint arXiv:2505.11063, 2025
2025 arXiv
-
[151]
ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Models
Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, and Min Yang. ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Models. arXiv preprint arXiv:2505.17244, 2025
2025
-
[152]
Deliberative alignment: Reasoning enables safer language models
Melody Y Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, et al. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339, 2024
2024 arXiv
-
[153]
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
Zijun Wang, Haoqin Tu, Yuhan Wang, Juncheng Wu, Jieru Mei, Brian R Bartoldson, Bhavya Kailkhura, and Cihang Xie. STAR-1: Safer Alignment of Reasoning LLMs with 1K Data. arXiv preprint arXiv:2504.01903, 2025
2025
-
[154]
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
Yichi Zhang, Zihao Zeng, Dongbai Li, Yao Huang, Zhijie Deng, and Yinpeng Dong. RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability. arXiv preprint arXiv:2504.10081, 2025
2025 arXiv
-
[155]
SAFEPATH: Preventing Harmful Reasoning in Chain-of- Thought via Early Alignment
Wonje Jeung, Sangyeon Yoon, Minsuk Kahng, and Albert No. SAFEPATH: Preventing Harmful Reasoning in Chain-of- Thought via Early Alignment. arXiv preprint arXiv:2505.14667, 2025
2025
-
[156]
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Heli Xu, Tianshu Chu, Peizhao Hu, and Yangqiu Song. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning. arXiv preprint arXiv:2505.14585, 2025
2025 arXiv
-
[157]
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation
Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. arXiv preprint arXiv:2503.11926, 2025
2025 arXiv
-
[158]
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, et al. How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study. arXiv preprint arXiv:2505.15404, 2025
2025 arXiv
-
[159]
Hair: Hardness-aware inverse reinforcement learning with introspective reasoning for llm alignment
Ruoxi Cheng, Haoxuan Ma, and Weixin Wang. Hair: Hardness-aware inverse reinforcement learning with introspective reasoning for llm alignment. arXiv preprint arXiv:2503.18991, 2025
2025
-
[160]
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, and Natasha Jaques. Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models. arXiv preprint arXiv:2506.07468, 2025
2025 arXiv
-
[161]
Safety tax: Safety alignment makes your large reasoning models less reasonable
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. Safety tax: Safety alignment makes your large reasoning models less reasonable. arXiv preprint arXiv:2503.00555, 2025
2025 arXiv
-
[162]
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
Kaiwen Zhou, Xuandong Zhao, Gaowen Liu, Jayanth Srinivasa, Aosong Feng, Dawn Song, and Xin Eric Wang. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning. arXiv preprint arXiv:2505.16186, 2025
2025
-
[163]
SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation
Naizhu Jin, Zhong Li, Yinggang Guo, Chao Su, Tian Zhang, and Qingkai Zeng. SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation. arXiv preprint arXiv:2412.05829, 2024
2024 arXiv
-
[164]
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
Zihao Zhu, Hongbao Zhang, Ruotong Wang, Ke Xu, Siwei Lyu, and Baoyuan Wu. To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models. arXiv preprint arXiv:2502.12202, 2025
2025 arXiv
-
[165]
Shadowcot: Cognitive hijacking for stealthy reasoning backdoors in llms
Gejian Zhao, Hanzhou Wu, Xinpeng Zhang, and Athanasios V Vasilakos. Shadowcot: Cognitive hijacking for stealthy reasoning backdoors in llms. arXiv preprint arXiv:2504.05605, 2025
2025 arXiv
-
[166]
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
James Chua, Jan Betley, Mia Taylor, and Owain Evans. Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models. arXiv preprint arXiv:2506.13206, 2025
2025 arXiv
-
[167]
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. Badchain: Backdoor chain-of-thought prompting for large language models. In Proc. ICLR, 2024
2024
-
[168]
Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun. Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models. arXiv preprint arXiv:2408.12798, 2024
2024 arXiv
-
[169]
Darkmind: Latent chain-of-thought backdoor in customized llms
Zhen Guo and Reza Tourani. Darkmind: Latent chain-of-thought backdoor in customized llms. arXiv preprint arXiv:2501.18617, 2025
2025
-
[170]
Process or result? manipulated ending tokens can mislead reasoning llms to ignore the correct reasoning steps
Yu Cui, Bryan Hooi, Yujun Cai, and Yiwei Wang. Process or result? manipulated ending tokens can mislead reasoning llms to ignore the correct reasoning steps. arXiv preprint arXiv:2503.19326, 2025
2025 arXiv
-
[171]
System prompt poisoning: Persistent attacks on large language models beyond user injection
Jiawei Guo and Haipeng Cai. System prompt poisoning: Persistent attacks on large language models beyond user injection. arXiv preprint arXiv:2505.06493, 2025
2025
-
[172]
Practical Reasoning Interruption Attacks on Reasoning Large Language Models
Yu Cui and Cong Zuo. Practical Reasoning Interruption Attacks on Reasoning Large Language Models. arXiv preprint arXiv:2505.06643, 2025. 30 References A Preprint
2025 arXiv
-
[173]
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
Yu Cui, Yujun Cai, and Yiwei Wang. Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression. arXiv preprint arXiv:2504.20493, 2025
2025 arXiv
-
[174]
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems
Hongru Song, Yu-an Liu, Ruqing Zhang, Jiafeng Guo, and Yixing Fan. Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems. arXiv preprint arXiv:2505.16367, 2025
2025 arXiv
-
[175]
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection
Ryan Marinelli, Josef Pichlmeier, and Tamas Bisztray. Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection. arXiv preprint arXiv:2503.21464, 2025
2025 arXiv
-
[176]
GUARD: Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generation
Naizhu Jin, Zhong Li, Tian Zhang, and Qingkai Zeng. GUARD: Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generation. arXiv preprint arXiv:2505.21425, 2025
2025 arXiv
-
[177]
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Xuandong Zhao, Wenxuan Zhang, Dawn Song, and Bingsheng He. Assessing Judging Bias in Large Reasoning Models: An Empirical Study. arXiv preprint arXiv:2504.09946, 2025
2025 arXiv
-
[178]
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
Wenxiao Wang, Parsa Hosseini, and Soheil Feizi. Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption. arXiv preprint arXiv:2504.20769, 2025
2025 arXiv
-
[179]
Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems? arXiv preprint arXiv:2504.00509, 2025
Kai Yan, Yufei Xu, Zhengyin Du, Xuesong Yao, Zheyu Wang, Xiaowen Guo, and Jiecao Chen. Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems? arXiv preprint arXiv:2504.00509, 2025
2025
-
[180]
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
Haoyan Yang, Runxue Bao, Cao Xiao, Jun Ma, Parminder Bhatia, Shangqian Gao, and Taha Kass-Hout. Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector. arXiv preprint arXiv:2505.17100, 2025
2025
-
[181]
Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models
Yuqing Wang and Yun Zhao. Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models. arXiv preprint arXiv:2406.11020, 2024
2024 arXiv
-
[182]
A Closer Look at System Prompt Robustness
Norman Mu, Jonathan Lu, Michael Lavery, and David Wagner. A Closer Look at System Prompt Robustness. arXiv preprint arXiv:2502.12197, 2025
2025 arXiv
-
[183]
A Frustratingly Simple Yet Highly Effective At- tack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.arXiv preprint arXiv:2503.10635, 2025
Zhaoyi Li, Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui, and Zhiqiang Shen. A Frustratingly Simple Yet Highly Effective At- tack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.arXiv preprint arXiv:2503.10635, 2025
2025
-
[184]
Reasoning Models Are More Easily Gaslighted Than You Think
Bin Zhu, Hailong Yin, Jingjing Chen, and Yu-Gang Jiang. Reasoning Models Are More Easily Gaslighted Than You Think. arXiv preprint arXiv:2506.09677, 2025
2025
-
[185]
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? In Proc
Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo, Zengmao Wang, and Bo Han. Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? In Proc. NeurIPS, 2024
2024
-
[186]
Stepwise Reasoning Disruption Attack of LLMs
Jingyu Peng, Maolin Wang, Xiangyu Zhao, Kai Zhang, Wanyu Wang, Pengyue Jia, Qidong Liu, Ruocheng Guo, and Qi Liu. Stepwise Reasoning Disruption Attack of LLMs. In Proc. ACL, pages 5040–5058, 2025
2025
-
[187]
Polymath: Evaluating mathematical reasoning in multilingual contexts
Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, et al. Polymath: Evaluating mathematical reasoning in multilingual contexts. arXiv preprint arXiv:2504.18428, 2025
2025
-
[188]
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
Meghana Rajeev, Rajkumar Ramamurthy, Prapti Trivedi, Vikas Yadav, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudan, James Zou, and Nazneen Rajani. Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models. arXiv preprint arXiv:2503.01781, 2025
2025 arXiv
-
[189]
Benchmarking reasoning robustness in large language models
Tong Yu, Yongcheng Jing, Xikun Zhang, Wentao Jiang, Wenjie Wu, Yingjie Wang, Wenbin Hu, Bo Du, and Dacheng Tao. Benchmarking reasoning robustness in large language models. arXiv preprint arXiv:2503.04550, 2025
2025 arXiv
-
[190]
MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations
Kaixuan Huang, Jiacheng Guo, Zihao Li, Xiang Ji, Jiawei Ge, Wenzhe Li, Yingqing Guo, Tianle Cai, Hui Yuan, Runzhe Wang, et al. MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations. arXiv preprint arXiv:2502.06453, 2025
2025 arXiv
-
[191]
CODECRASH: Stress Testing LLM Reasoning under Structural and Semantic Perturbations
Man Ho Lam, Chaozheng Wang, Jen-tse Huang, and Michael R Lyu. CODECRASH: Stress Testing LLM Reasoning under Structural and Semantic Perturbations. arXiv preprint arXiv:2504.14119, 2025
2025
-
[192]
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
Jaechul Roh, Varun Gandhi, Shivani Anilkumar, and Arin Garg. Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation. arXiv preprint arXiv:2506.06971, 2025
2025 arXiv
-
[193]
Preemptive answer “attacks" on chain-of-thought reasoning
Rongwu Xu, Zehan Qi, and Wei Xu. Preemptive answer “attacks" on chain-of-thought reasoning. In Findings of Proc. ACL, 2024
2024
-
[194]
Large language models are unconscious of unreasonability in math problems
Jingyuan Ma, Damai Dai, Lei Sha, and Zhifang Sui. Large language models are unconscious of unreasonability in math problems. arXiv preprint arXiv:2403.19346, 2024
2024 arXiv
-
[195]
Dnr bench: Benchmarking over-reasoning in reasoning llms
Masoud Hashemi, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudhan, Jishnu Sethumadhavan Nair, Aman Tiwari, and Vikas Yadav. Dnr bench: Benchmarking over-reasoning in reasoning llms. arXiv preprint arXiv:2503.15793, 2025
2025 arXiv
-
[196]
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? In Proc
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, et al. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? In Proc. ACL, 2025
2025
-
[197]
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? arXiv preprint arXiv:2504.06514, 2025
Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou. Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? arXiv preprint arXiv:2504.06514, 2025. 31 References A Preprint
2025 arXiv
-
[198]
Excessive Reasoning Attack on Reasoning LLMs
Wai Man Si, Mingjie Li, Michael Backes, and Yang Zhang. Excessive Reasoning Attack on Reasoning LLMs. arXiv preprint arXiv:2506.14374, 2025
2025 arXiv
-
[199]
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, et al. Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs. arXiv preprint arXiv:2501.18585, 2025
2025 arXiv
-
[200]
Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms
Jinyan Su, Jennifer Healey, Preslav Nakov, and Claire Cardie. Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms. arXiv preprint arXiv:2505.00127, 2025
2025 arXiv
-
[201]
Internal Bias in Reasoning Models leads to Overthinking
Renfei Dang, Shujian Huang, and Jiajun Chen. Internal Bias in Reasoning Models leads to Overthinking. arXiv preprint arXiv:2505.16448, 2025
2025
-
[202]
Overthink: Slowdown attacks on reasoning llms
Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, and Eugene Bagdasarian. Overthink: Slowdown attacks on reasoning llms. arXiv preprint arXiv:2502.02542, 2025
2025
-
[203]
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al. The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks. arXiv preprint arXiv:2502.08235, 2025
2025 arXiv
-
[204]
Output Length Effect on DeepSeek-R1’s Safety in Forced Thinking
Xuying Li, Zhuo Li, Yuji Kosuga, and Victor Bian. Output Length Effect on DeepSeek-R1’s Safety in Forced Thinking. arXiv preprint arXiv:2503.01923, 2025
2025 arXiv
-
[205]
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
Chung-En Sun, Ge Yan, and Tsui-Wei Weng. ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models. arXiv preprint arXiv:2503.22048, 2025
2025
-
[206]
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J Wooldridge, Janet B Pierrehumbert, and Furu Wei. Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks. In Proc. ACL, 2025
2025
-
[207]
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
Xiaoqing Cheng, Hongying Zan, Lulu Kong, Jinwang Song, and Min Peng. Detection, Classification, and Mitigation of Gender Bias in Large Language Models. arXiv preprint arXiv:2506.12527, 2025
2025 arXiv
-
[208]
Prompting techniques for reducing social bias in llms through system 1 and system 2 cognitive processes
Mahammed Kamruzzaman and Gene Louis Kim. Prompting techniques for reducing social bias in llms through system 1 and system 2 cognitive processes. arXiv preprint arXiv:2404.17218, 2024
2024 arXiv
-
[209]
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
Saloni Dash, Amélie Reymond, Emma S Spiro, and Aylin Caliskan. Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning. arXiv preprint arXiv:2506.20020, 2025
2025 arXiv
-
[210]
Bias runs deep: Implicit reasoning biases in persona-assigned llms
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: Implicit reasoning biases in persona-assigned llms. In Proc. ICLR, 2024
2024
-
[211]
Biasguard: A reasoning-enhanced bias detection tool for large language models
Zhiting Fan, Ruizhe Chen, and Zuozhu Liu. Biasguard: A reasoning-enhanced bias detection tool for large language models. In Findings of Proc. ACL, 2025
2025
-
[212]
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
Riccardo Cantini, Nicola Gabriele, Alessio Orsino, and Domenico Talia. Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models. arXiv preprint arXiv:2507.02799, 2025
2025 arXiv
-
[213]
R-tofu: Unlearning in large reasoning models
Sangyeon Yoon, Wonje Jeung, and Albert No. R-tofu: Unlearning in large reasoning models. arXiv preprint arXiv:2505.15214, 2025
2025 arXiv
-
[214]
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
Changsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, and Sijia Liu. Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills. arXiv preprint arXiv:2506.12963, 2025
2025
-
[215]
Step-by-Step Reasoning Attack: Revealing ’Erased’ Knowledge in Large Language Models
Yash Sinha, Manit Baser, Murari Mandal, Dinil Mon Divakaran, and Mohan Kankanhalli. Step-by-Step Reasoning Attack: Revealing ’Erased’ Knowledge in Large Language Models. arXiv preprint arXiv:2506.17279, 2025
2025 arXiv
-
[216]
ImF: Implicit Fingerprint for Large Language Models
Peng Wanli, Xue Yiming, et al. ImF: Implicit Fingerprint for Large Language Models. arXiv preprint arXiv:2503.21805, 2025
2025 arXiv
-
[217]
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
Zhenzhen Ren, GuoBiao Li, Sheng Li, Zhenxing Qian, and Xinpeng Zhang. CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models. arXiv preprint arXiv:2505.16785, 2025
2025 arXiv
-
[218]
Towards copyright protection for knowledge bases of retrieval-augmented language models via ownership verification with reasoning
Junfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu, Chenxi Liu, Yanshuo Chen, and Heng Huang. Towards copyright protection for knowledge bases of retrieval-augmented language models via ownership verification with reasoning. arXiv preprint arXiv:2502.10440, 2025
2025 arXiv
-
[219]
Antidistillation sampling
Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild, Alexander Robey, Marc Finzi, and J Zico Kolter. Antidistillation sampling. arXiv preprint arXiv:2504.13146, 2025
2025
-
[220]
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
Tommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun, and Seong Joon Oh. Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers. arXiv preprint arXiv:2506.15674, 2025
2025
-
[221]
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
Weidi Luo, Tianyu Lu, Qiming Zhang, Xiaogeng Liu, Bin Hu, Yue Zhao, Jieyu Zhao, Song Gao, Patrick McDaniel, Zhen Xiang, et al. Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models. arXiv preprint arXiv:2504.19373, 2025
2025
-
[222]
TrustLLM: Trustworthiness in Large Language Models
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, et al. TrustLLM: Trustworthiness in Large Language Models. In Proc. ICML, 2024. 32 References A Preprint
2024
-
[223]
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Informatio...
2025
-
[224]
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das. A survey of hallucination in large foundation models. arXiv preprint arXiv:2309.05922, 2023
2023 arXiv
-
[225]
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation.arXiv preprint arXiv:2506.17088, 2025
Jiahao Cheng, Tiancheng Su, Jia Yuan, Guoxiu He, Jiawei Liu, Xinqi Tao, Jingwen Xie, and Huaxia Li. Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation.arXiv preprint arXiv:2506.17088, 2025
2025
-
[226]
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024
2024 arXiv
-
[227]
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In Proc. ICLR, 2023
2023
-
[228]
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Proc. ACL, 2022
2022
-
[229]
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. Halueval: A large-scale hallucination evaluation benchmark for large language models. In Proc. EMNLP, 2023
2023
-
[230]
Evaluating hallucinations in chinese large language models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, et al. Evaluating hallucinations in chinese large language models. arXiv preprint arXiv:2310.03368, 2023
2023 arXiv
-
[231]
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368, 2024
2024 arXiv
-
[232]
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017
2017 arXiv
-
[233]
Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods
Wei Li, Wenhao Wu, Moye Chen, Jiachen Liu, Xinyan Xiao, and Hua Wu. Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods. arXiv preprint arXiv:2203.05227, 2022
2022 arXiv
-
[234]
Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? In Proc
Alon Jacovi and Yoav Goldberg. Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? In Proc. ACL, pages 4198–4205, 2020
2020
-
[235]
Dissociation of faithful and unfaithful reasoning in llms
Evelyn Yee, Alice Li, Chenyu Tang, Yeon Ho Jung, Ramamohan Paturi, and Leon Bergen. Dissociation of faithful and unfaithful reasoning in llms. arXiv preprint arXiv:2405.15092, 2024
2024 arXiv
-
[236]
Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language? In Findings of Proc
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language? In Findings of Proc. EMNLP, pages 4351–4367, 2020
2020
-
[237]
Negative preference optimization: From catastrophic collapse to effective unlearning
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei. Negative preference optimization: From catastrophic collapse to effective unlearning. In Proc. COLM, 2024
2024
-
[238]
Causality
Judea Pearl. Causality. Cambridge university press, 2009
2009
-
[239]
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Findings of Proc. ACL, p...
2023
-
[240]
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. In Proc. ICML, 2024
2024
-
[241]
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al. A strongreject for empty jailbreaks. In Proc. NeurIPS D&B Track, 2024
2024
-
[242]
Air-bench 2024: A safety benchmark based on risk categories from regulations and policies
Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, et al. Air-bench 2024: A safety benchmark based on risk categories from regulations and policies. arXiv preprint arXiv:2407.17436, 2024
2024 arXiv
-
[243]
WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs. In Proc. NeurIPS D&B Track, 2024
2024
-
[244]
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023
2023 arXiv
-
[245]
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In Proc. SaTML, 2025
2025
-
[246]
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. In Proc. NeurIPS, 2024
2024
-
[247]
Gemini 2.0 Flash Thinking, 2025
Google DeepMind. Gemini 2.0 Flash Thinking, 2025. 33 References A Preprint
2025
-
[248]
Kimi k1.5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025
2025 arXiv
-
[249]
Sky-T1: Train your own O1 preview model within $450
NovaSky Team. Sky-T1: Train your own O1 preview model within $450. https://novasky-ai.github.io/posts/sky-t1, 2025. Accessed: 2025-01-09
2025
-
[250]
QwQ-32B: Embracing the Power of Reinforcement Learning, March 2025
Qwen Team. QwQ-32B: Embracing the Power of Reinforcement Learning, March 2025
2025
-
[251]
Skywork-o1 Open Series
Skywork o1 Team. Skywork-o1 Open Series. https://huggingface.co/Skywork, November 2024
2024
-
[252]
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models. In Proc. NeurIPS, 2024
2024
-
[253]
Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models
Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, et al. Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models. arXiv...
2024 arXiv
-
[254]
Chisafety- bench: A chinese hierarchical safety benchmark for large language models
Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An, Bikun Yang, KaiKai Zhao, Kai Wang, and Shiguo Lian. Chisafety- bench: A chinese hierarchical safety benchmark for large language models. arXiv preprint arXiv:2406.10311, 2024
2024 arXiv
-
[255]
Jailbroken: How does llm safety training fail? In Proc
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? In Proc. NeurIPS, 2023
2023
-
[256]
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, et al. R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization. arXiv preprint arXiv:2503.10615, 2025
2025 arXiv
-
[257]
Skywork r1v: Pioneering multimodal reasoning with chain-of-thought
Yi Peng, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, et al. Skywork r1v: Pioneering multimodal reasoning with chain-of-thought. arXiv preprint arXiv:2504.05599, 2025
2025 arXiv
-
[258]
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, and Xu Yang. Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl. arXiv preprint arXiv:2503.07536, 2025
2025 arXiv
-
[259]
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation. In Proc. ECCV, pages 388–404, 2024
2024
-
[260]
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models? arXiv preprint arXiv:2504.10000, 2025
Yanbo Wang, Jiyang Guan, Jian Liang, and Ran He. Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models? arXiv preprint arXiv:2504.10000, 2025
2025 arXiv
-
[261]
More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models. arXiv preprint arXiv:2302.12173, 2023
2023 arXiv
-
[262]
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. Defending chatgpt against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12):1486–1496, 2023
2023
-
[263]
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, and Minlie Huang. Defending large language models against jailbreaking attacks through goal prioritization. In Proc. ACL, 2024
2024
-
[264]
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387, 2023
2023 arXiv
-
[265]
Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho. Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks. In Findings of Proc. ACL, 2025
2025
-
[266]
Root defence strategies: Ensuring safety of llm at the decoding level
Xinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang, and Yu Tian. Root defence strategies: Ensuring safety of llm at the decoding level. arXiv preprint arXiv:2410.06809, 2024
2024 arXiv
-
[267]
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. Safedecoding: Defending against jailbreak attacks via safety-aware decoding. arXiv preprint arXiv:2402.08983, 2024
2024 arXiv
-
[268]
Safeinfer: Context adaptive decoding time safety alignment for large language models
Somnath Banerjee, Sayan Layek, Soham Tripathy, Shanu Kumar, Animesh Mukherjee, and Rima Hazra. Safeinfer: Context adaptive decoding time safety alignment for large language models. In Proc. AAAI, pages 27188–27196, 2025
2025
-
[269]
Safeguarding large language models: A survey
Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, et al. Safeguarding large language models: A survey. arXiv preprint arXiv:2406.02622, 2024
2024 arXiv
-
[270]
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023
2023 arXiv
-
[271]
Llama guard 3 vision: Safeguarding human-ai image understanding conversations
Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti. Llama guard 3 vision: Safeguarding human-ai image understanding conversations. arXiv preprint arXiv:2411.10414, 2024
2024 arXiv
-
[272]
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
llama Team. Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. https://ai.meta.com/blog/llama- 3-2-connect-2024-vision-edge-mobile-devices/, 2024. Accessed: 2024-09-25. 34 References A Preprint
2024
-
[273]
A comprehensive survey of LLM alignment techniques: RLHF, RLAIF, PPO, DPO and more
Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al. A comprehensive survey of LLM alignment techniques: RLHF, RLAIF, PPO, DPO and more. arXiv preprint arXiv:2407.16216, 2024
2024 arXiv
-
[274]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In Proc. NeurIPS, 2022
2022
-
[275]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022
2022 arXiv
-
[276]
Rlaif vs
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, et al. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. In Proc. ICML, 2024
2024
-
[277]
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. In Proc. NeurIPS, 2024
2024
-
[278]
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference. In Proc. ACL, 2025
2025
-
[279]
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. In Proc. ICML, 2024
2024
-
[280]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. In Proc. ICLR, 2022
2022
-
[281]
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Proc. ...
2024
-
[282]
STAIR: Improving Safety Alignment with Introspective Reasoning
Yichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia, Zhengwei Fang, Xiao Yang, Ranjie Duan, Dong Yan, Yinpeng Dong, and Jun Zhu. STAIR: Improving Safety Alignment with Introspective Reasoning. arXiv preprint arXiv:2502.02384, 2025
2025 arXiv
-
[283]
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. In Findings of Proc. ACL, 2024
2024
-
[284]
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707, 2023
2023 arXiv
-
[285]
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. In Findings of Proc. EMNLP, 2023
2023
-
[286]
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al. Sorry-bench: Systematically evaluating large language model safety refusal behaviors. In Proc. ICLR, 2025
2025
-
[287]
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. In Proc. NAACL, 2024
2024
-
[288]
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks. In Proc. COLM, 2024
2024
-
[289]
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. In Proc. NeurIPS, 2023
2023
-
[290]
Simplesafetytests: a test suite for identifying critical safety risks in large language models
Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk, Rebecca Qian, Anand Kannappan, Scott A Hale, and Paul Röttger. Simplesafetytests: a test suite for identifying critical safety risks in large language models. arXiv preprint arXiv:2311.08370, 2023
2023 arXiv
-
[291]
Tdc 2023 (llm edition): The trojan detection challenge
Mazeika Mantas, Zou Andy, Mu Norman, Phan Long, Wang Zifan, Yu Chunru, Khoja Adam, Jiang Fengqing, O’Gara Aidan, Sakhaee Ellie, Xiang Zhen, Rajabi Arezoo, Hendrycks Dan, Poovendran Radha, Li Bo, and Forsyth David. Tdc 2023 (llm edition): The trojan detection challenge. In Proc...
2023
-
[292]
ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming
Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li. ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming. arXiv preprint arXiv:2404.08676, 2024
2024 arXiv
-
[293]
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, et al. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proc. CVPR, pages 13807–13816, 2024
2024
-
[294]
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. In Findings of Proc. ACL, 2024
2024
-
[295]
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, and Qi Liu. VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment. In Proc. EMNLP, 2024. 35 References A Preprint
2024
-
[296]
Safe RLHF-V: Safe Reinforcement Learning from Human Feedback in Multimodal Large Language Models
Jiaming Ji, Xinyu Chen, Rui Pan, Han Zhu, Conghui Zhang, Jiahao Li, Donghai Hong, Boyuan Chen, Jiayi Zhou, Kaile Wang, et al. Safe RLHF-V: Safe Reinforcement Learning from Human Feedback in Multimodal Large Language Models. arXiv preprint arXiv:2503.17682, 2025
2025 arXiv
-
[297]
Mm-rlhf: The next step forward in multimodal llm alignment
Yi-Fan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu, Peiyan Li, Jianshu Zeng, Wulin Xie, Yang Shi, Huanyu Zhang, Junkang Wu, et al. Mm-rlhf: The next step forward in multimodal llm alignment. In Proc. ICML, 2025
2025
-
[298]
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014
2014 arXiv
-
[299]
Mitigating the alignment tax of rlhf
Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, et al. Mitigating the alignment tax of rlhf. In Proc. EMNLP, 2024
2024
-
[300]
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the ninth workshop on statis...
2014
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.