Pith. sign in

REVIEW 3 major objections 6 minor 8 cited by

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Long-CoT reasoning is no trustworthiness upgrade, survey finds.

desk verdict Useful map of reasoning trustworthiness literature, but the abstract's 'comparable or even greater' claim is undercut by the body's own benchmark-dependent evidence; revise, don't reject. read the letter →

arxiv 2509.03871 v1 pith:P5SZ5446 submitted 2025-09-04 cs.CL cs.AIcs.CR

classification cs.CLcs.AIcs.CR
keywords chain-of-thoughtreasoninglargemodelstrustworthinessjailbreaksafetyhallucinationfaithfulnessrobustnessprivacyleakage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey asks what long chain-of-thought reasoning does to trustworthiness across five axes: truthfulness, safety, robustness, fairness, and privacy. It argues that reasoning techniques genuinely help in some areas, such as hallucination detection and harmful-content screening, while current reasoning models simultaneously show comparable or larger vulnerabilities than ordinary chat models on safety, robustness, and privacy. The recurring object is the thinking trace itself: it can be attacked, it can be unfaithful, it leaks private attributes, and it can make harmful answers more detailed and persuasive. The survey's organizing conclusion is that reasoning capability is not an automatic trustworthiness upgrade; it shifts where failures occur and expands the attack surface.

What carries the argument

The key machinery is the survey's taxonomy of trustworthy reasoning: five dimensions (truthfulness, safety, robustness, fairness, privacy) crossed with two implementation paradigms (chain-of-thought prompting and end-to-end large reasoning models). Within that grid, the load-bearing object is the chain of thought itself—the generated intermediate reasoning steps that serve simultaneously as a faithfulness measurement target, a jailbreak surface, a backdoor trigger channel, a source of overthinking and underthinking, and a privacy side channel.

What would settle it

Run a standardized safety battery on the same reasoning models across many languages and attack templates and check whether open-source reasoning models remain consistently less safe than their base counterparts; if the rank order reverses across datasets, the pooled conclusion about comparable or greater vulnerabilities weakens.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that while chain-of-thought prompting and reasoning models can improve truthfulness through hallucination mitigation and support safety defenses such as guardrail models, the same reasoning capability produces new failure modes: reasoning models hallucinate in simple non-reasoning tasks, generate more harmful content after jailbreak, are less robust to small input perturbations, preserve or amplify bias, and leak more private information through their thinking traces. The synthesis, stated in the abstract, is that current reasoning models 'often suffer from comparable or even greater vulnerabilities' in safety, robustness, and privacy, with the

Load-bearing premise

The survey's cross-model conclusions assume that safety and robustness results measured on different benchmarks, languages, and attack templates can be combined into a single verdict; the paper itself reports cases where model rankings flip depending on the dataset.

Editorial extensions

If this is right

  • Safety auditing of deployed reasoning models should monitor the thinking trace, not just the final answer, because several cited studies find the reasoning content is less safe than the output.
  • Jailbreak defenses validated on chat models cannot be assumed to transfer to reasoning models; new attacks in this taxonomy specifically exploit reasoning traces, ciphers, or stepwise decomposition.
  • Faithfulness evaluation needs standardized protocols; current intervention metrics can confound model strength with apparent unfaithfulness.
  • Aligning reasoning models requires chain-of-thought-specific data and training, since safety reasoning must be taught rather than assumed to follow from general reasoning ability.
  • Privacy and interpretability conflict: visible thinking traces improve transparency but also enable attribute inference and unlearning-recovery attacks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not claimed by the paper, but the dataset-dependence of safety rankings suggests that single-number attack-success-rate scores can mislead; reporting per-category and per-language results would be a cheap test of which vulnerabilities are stable.
  • Not claimed by the paper, but the finding that shortened reasoning improves harmlessness points to a concrete deployable intervention: force short reasoning at inference time and measure safety across languages.
  • Not claimed by the paper, but if the safety tax scales with reasoning-length incentives, separating rewards for correctness from rewards for safety during reinforcement learning may decouple the trade-off.
  • Not claimed by the paper, but if the multilingual vulnerabilities are a form of mismatched generalization, then training and evaluating safety in a few dominant languages systematically underestimates real-world deployment risk.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey reviews the trustworthiness of chain-of-thought prompting and large reasoning models (LRMs) across five dimensions: truthfulness, safety, robustness, fairness, and privacy. It organizes recent work into a taxonomy, summarizes methods and findings in each area, and identifies open problems. The central claim, stated in the abstract and Introduction, is that reasoning techniques can improve some aspects of trustworthiness (e.g., hallucination mitigation, harmful-content detection, robustness) while cutting-edge reasoning models themselves often exhibit comparable or greater vulnerabilities in safety, robustness, and privacy than non-reasoning models. The paper positions itself as the first comprehensive survey of trustworthy reasoning and provides a public repository of related papers.

Significance. If the synthesis is reliable, the survey would be a useful resource for the AI-safety community. Its main strengths are the breadth of covered topics (five trustworthiness dimensions), the explicit separation of early CoT techniques from end-to-end reasoning models, and the identification of several concrete open problems, such as the need for standardized faithfulness metrics and more fine-grained benchmarks. The paper also highlights a genuinely important tension: improved reasoning capability does not automatically translate into improved trustworthiness. The structured taxonomy and the GitHub resource add practical value. However, the central comparative claim—that reasoning models 'often suffer from comparable or even greater vulnerabilities'—is only partially supported by the evidence the paper itself presents, because the underlying evaluations are heterogeneous and sometimes contradictory. The survey also contains several citation and attribution errors that undermine reliability, especially for a resource whose purpose is to guide readers to the correct literature.

major comments (3)
  1. [Abstract and §4.1] The abstract claims that reasoning models 'often suffer from comparable or even greater vulnerabilities' in safety, robustness, and privacy. This comparative claim is not stable under the evidence reported in §4.1. The paper's own third bullet explicitly states that 'pairwise safety ranks between models depend on datasets': AirBench finds DeepSeek-R1 safer than DeepSeek-V3, while CNSafe finds the opposite with an average ASR margin of 21.7%, and WildGuard Jailbreak vs CNSafe_RT again reverse the picture. Similar benchmark-dependence appears in robustness (§5.2: reasoning models beat non-reasoning models on CodeCrash but lose on Math-RoB and other perturbation benchmarks). Since these evaluations use different attack templates, languages, and model versions, pooling them into a single 'comparable or even greater' verdict is unsupported. The abstract and conclusion should either be qualifi
  2. [Introduction and §2.2] Several references are mischaracterized or misassigned. In the Introduction, reference [12] is described as 'a related survey [12] provided valuable discussions on safety-related aspects,' but [12] is 'Safety Reasoning with Guidelines' (ICML 2025), an original research paper, not a survey. In §2.2, the sentence 'zero-shot-CoT [16]' attributes zero-shot CoT to Wei et al. [16]; the correct reference for zero-shot CoT is Kojima et al. [17]. In addition, 'few-shot-CoT [19]' in the same paragraph is wrong: [19] is the Chain-of-Scrutiny backdoor-detection paper, not the few-shot CoT paper. For a survey whose purpose is to organize and signpost the literature, these citation errors are material and should be corrected throughout.
  3. [§3.2.2 and §8] The faithfulness section reports directly contradictory conclusions (e.g., 'larger models are generally more faithful' [85,78] vs 'models with higher accuracy tend to exhibit lower faithfulness' [79,86]; 'length penalties may result in unfaithful responses' [81] vs 'unfaithful CoTs are usually longer' [82]). The survey notes these contradictions and calls for standardized metrics, which is appropriate. However, the synthesis does not identify which differences are likely due to evaluation methodology (e.g., intervention type, task difficulty, model family) versus genuine model properties. Since the paper explicitly lists 'standard measurements of faithfulness' as a future direction, it should at least organize the existing evidence around the methodological axes that are already discussed in §3.2.1. Without that, the 'factors that influence faithfulness' subsection is more a list of conf
minor comments (6)
  1. [§2.2] The citation typo 'zero-shot-CoT [16]' should read [17]; the nearby 'few-shot-CoT [19]' should read [16] (or the intended reference). Please check all citation numbers against the bibliography.
  2. [§3.2.1] Typo: 'Lakage-Adjusted Simulatability' should be 'Leakage-Adjusted Simulatability'; 'Paulet al.' should be 'Paul et al.'
  3. [§4.3.3] Typo: 'convolution neural network' should be 'convolutional neural network.'
  4. [Figure 2] The taxonomy figure is dense and the font is small. Consider grouping by sub-theme or providing an accompanying table to improve readability.
  5. [References] Reference formats are inconsistent: some entries include only arXiv identifiers and no publication venue, while others include venue names. For a survey, adding DOIs or stable URLs would improve usability.
  6. [§6] The fairness section is comparatively short and does not discuss the interaction of fairness with reasoning-model training (e.g., RLVR) beyond citing a few benchmarks. A brief discussion of open problems in this specific intersection would strengthen the survey.

Circularity Check

0 steps flagged · score 2.0 of 10

No material circularity: survey synthesis rests on external literature; self-citations are minor and not load-bearing.

full rationale

This is a survey paper rather than a derivation; its central claim—that reasoning techniques can improve some aspects of trustworthiness while reasoning models themselves show comparable or greater vulnerabilities—is a synthesis of many independent empirical studies, not the output of a fitted model or a self-referential equation. No step was found where a parameter is fit to a subset of data and then renamed a prediction, and no uniqueness theorem from the authors' own prior work is invoked to force a conclusion. The authors do cite several of their own papers ([260], [304], [305], [340], [344], [345]), but in all cases these are background or supporting citations within broader enumerations of related work; the central conclusions are supported by numerous third-party benchmarks and studies. The paper even explicitly acknowledges benchmark-dependent disagreements in Section 4.1 ('Pairwise safety ranks between models depend on datasets'), which is a limitation on comparability rather than a circularity. Under the hard rules requiring a specific quoted reduction for a circularity finding, no such reduction can be exhibited. The score of 2 reflects the presence of minor self-citations that are not load-bearing, not any actual circular reasoning.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey is a literature synthesis, not a derivation, so it has no free parameters or invented entities. Its claims rest on the adequacy of its taxonomy, the reliability of the cited evaluations, and the comparability of heterogeneous benchmarks.

assumptions (3)
  • domain assumption The five categories (truthfulness, safety, robustness, fairness, privacy) constitute an adequate decomposition of trustworthiness.
    Section 1 and Figure 2 organize the entire survey around these five dimensions; the paper does not justify exhaustiveness or test alternative taxonomies.
  • domain assumption The summarized findings in the cited papers are reliable and representative of the literature up to June 2025.
    The survey does not re-run evaluations; its synthesis about reasoning models having comparable or greater vulnerabilities depends on trusting the cited benchmarks.
  • domain assumption Conflicting evaluation results across benchmarks can be reconciled into a cross-model conclusion.
    Section 4.1 states that pairwise safety ranks depend on datasets (Airbench vs CNSafe), yet the paper still concludes that reasoning models are vulnerable; the reconciliation rule is not stated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models." pith.science (2026). https://pith.science/paper/P5SZ5446

@misc{pith2026250903871,
  author       = {Pith},
  title        = {Pith review of: A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P5SZ5446}},
  note         = {Machine review of arXiv:2509.03871}
}
read the original abstract

The development of Long-CoT reasoning has advanced LLM performance across various tasks, including language understanding, complex problem solving, and code generation. This paradigm enables models to generate intermediate reasoning steps, thereby improving both accuracy and interpretability. However, despite these advancements, a comprehensive understanding of how CoT-based reasoning affects the trustworthiness of language models remains underdeveloped. In this paper, we survey recent work on reasoning models and CoT techniques, focusing on five core dimensions of trustworthy reasoning: truthfulness, safety, robustness, fairness, and privacy. For each aspect, we provide a clear and structured overview of recent studies in chronological order, along with detailed analyses of their methodologies, findings, and limitations. Future research directions are also appended at the end for reference and discussion. Overall, while reasoning techniques hold promise for enhancing model trustworthiness through hallucination mitigation, harmful content detection, and robustness improvement, cutting-edge reasoning models themselves often suffer from comparable or even greater vulnerabilities in safety, robustness, and privacy. By synthesizing these insights, we hope this work serves as a valuable and timely resource for the AI safety community to stay informed on the latest progress in reasoning trustworthiness. A full list of related papers can be found at \href{https://github.com/ybwang119/Awesome-reasoning-safety}{https://github.com/ybwang119/Awesome-reasoning-safety}.

Figures

Figures reproduced from arXiv: 2509.03871 by the authors.

Figure 1
Figure 1. Illustration of typical CoT prompting. Few-shot-CoT uses several examples with the reasoning process to elicit CoT, and [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of trustworthiness in reasoning with large language models. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. An example of H-CoT jailbreak prompt, which is from “DukeCEICenter/Malicious_Educator_hcot_o1” dataset [117]. [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Examples of prompts for CoT data synthesis. Minor modifications are executed for better readability. [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Risky Business: Measuring The Faithfulness-Safety Tension

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Faithful reasoning and safety pull in opposite directions in current reasoning models, and the two behaviors are controlled by anti-correlated internal vectors that can be steered independently.

  2. Where Do CoT Training Gains Land in LLM based Agents?

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    CoT training in LLM agents improves prompt-action quality more than the advantage of generated reasoning, and selectively masking action supervision improves out-of-domain generalization.

  3. Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries

    cs.LG 2026-05 conditional novelty 6.0 of 10

    Swapping the reasoning trace prefill on unlearned weights can replicate or reverse the parser-split bypass gap, showing that the gap alone does not identify or rule out weight-level memorization.

  4. Pause or Fabricate? Training Language Models for Grounded Reasoning

    cs.CL 2026-04 conditional novelty 6.0 of 10

    GRIL uses stage-specific RL rewards to train LLMs to detect missing premises, pause proactively, and resume grounded reasoning after clarification, yielding up to 45% better premise detection and 30% higher task succe...

  5. From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    PreRL applies reward-driven updates to P(y) in pre-train space, uses Negative Sample Reinforcement to prune bad reasoning paths and boost reflection, and combines with standard RL in Dual Space RL to outperform baseli...

  6. Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

    cs.CR 2026-02 conditional novelty 6.0 of 10

    TRACE-RPS drops LLM attribute inference accuracy from around 50% to below 5% via fine-grained anonymization plus a two-stage rejection optimization.

  7. Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    Faithful Warm-Start pre-training on causally consistent vision-language samples improves accuracy, stabilizes RL, and reduces unsupported reasoning in VLMs.

  8. Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework

    cs.CR 2026-04 unverdicted novelty 5.0 of 10

    A 16-factor structured prompt framework strengthens CoT reasoning in LLMs for security analysis, yielding up to 40% reasoning gains in smaller models and stable accuracy improvements validated by human raters with Coh...

Reference graph

Works this paper leans on

299 extracted references · 6 canonical work pages · cited by 8 Pith papers

  1. [12]

    Safety Reasoning with Guidelines

    Haoyu Wang, Zeyu Qin, Li Shen, Xueqian Wang, Dacheng Tao, and Minhao Cheng. Safety Reasoning with Guidelines. In Proc. ICML, 2025

  2. [16]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. In Proc. NeurIPS, 2022

  3. [17]

    Large language models are zero-shot reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Proc. NeurIPS, 2022

  4. [19]

    Chain-of-scrutiny: Detecting backdoor attacks for large language models

    Xi Li, Yusen Zhang, Renze Lou, Chen Wu, and Jiaqi Wang. Chain-of-scrutiny: Detecting backdoor attacks for large language models. arXiv preprint arXiv:2406.05948, 2024

  5. [81]

    Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025

    James Chua and Owain Evans. Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025

  6. [82]

    Reasoning Models Don’t Always Say What They Think

    Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner Fabien Roger Vlad Mikulik, Sam Bowman, Jan Leike Jared Kaplan, et al. Reasoning Models Don’t Always Say What They Think. Anthropic Research, 2025

  7. [1]

    Safechain: Safety of language models with long chain-of-thought reasoning capabilities

    Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. Safechain: Safety of language models with long chain-of-thought reasoning capabilities. arXiv preprint arXiv:2502.12025, 2025

  8. [2]

    Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking? arXiv preprint arXiv:2505.17650, 2025

    Chengda Lu, Xiaoyu Fan, Yu Huang, Rongwu Xu, Jijie Li, and Wei Xu. Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking? arXiv preprint arXiv:2505.17650, 2025

Show all 299 references
  1. [3]

    Towards understanding the safety boundaries of deepseek models: Evaluation and findings

    Zonghao Ying, Guangyi Zheng, Yongxin Huang, Deyue Zhang, Wenxin Zhang, Quanchen Zou, Aishan Liu, Xianglong Liu, and Dacheng Tao. Towards understanding the safety boundaries of deepseek models: Evaluation and findings. arXiv preprint arXiv:2503.15092, 2025

  2. [4]

    A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment

    Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, et al. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585, 2025

  3. [5]

    Attacks, defenses and evaluations for llm conversation safety: A survey

    Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. Attacks, defenses and evaluations for llm conversation safety: A survey. In Proc. NAACL, 2024

  4. [6]

    Large language model safety: A holistic survey

    Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li, Yongqi Leng, Renren Jin, Chuang Liu, Xinwei Wu, Zishan Guo, Linhao Yu, et al. Large language model safety: A holistic survey. arXiv preprint arXiv:2412.17686, 2024

  5. [7]

    Towards reasoning era: A survey of long chain-of-thought for reasoning large language models

    Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wangxiang Che. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models. arXiv preprint arXiv:2503.09567, 2025

  6. [8]

    Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

    Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al. Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models. arXiv preprint arXiv:2501.09686, 2025

  7. [9]

    A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond

    Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, et al. A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond. arXiv preprint arXiv:2503.21614, 2025

  8. [10]

    Stop overthinking: A survey on efficient reasoning for large language models

    Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, et al. Stop overthinking: A survey on efficient reasoning for large language models. arXiv preprint arXiv:2503.16419, 2025

  9. [11]

    Efficient reasoning models: A survey

    Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. Efficient reasoning models: A survey. arXiv preprint arXiv:2504.10903, 2025

  10. [13]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  11. [14]

    The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

    Anthropic. The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

  12. [15]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  13. [18]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In Proc. NeurIPS, 2020

  14. [20]

    Openai o1 system card

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  15. [21]

    Openr: An open source framework for advanced reasoning with large language models

    Jun Wang, Meng Fang, Ziyu Wan, Muning Wen, Jiachen Zhu, Anjie Liu, Ziqin Gong, Yan Song, Lei Chen, Lionel M Ni, et al. Openr: An open source framework for advanced reasoning with large language models. arXiv preprint arXiv:2410.09671, 2024

  16. [22]

    O1 Replication Journey: A Strategic Progress Report–Part 1

    Yiwei Qin, Xuefeng Li, Haoyang Zou, Yixiu Liu, Shijie Xia, Zhen Huang, Yixin Ye, Weizhe Yuan, Hector Liu, Yuanzhi Li, et al. O1 Replication Journey: A Strategic Progress Report–Part 1. arXiv preprint arXiv:2410.18982, 2024

  17. [23]

    O1 Replication Journey–Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? arXiv preprint arXiv:2411.16489, 2024

    Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu. O1 Replication Journey–Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson? arXiv preprint arXiv:2411.16489, 20...

  18. [24]

    O1 Replication Journey–Part 3: Inference-time Scaling for Medical Reasoning

    Zhongzhen Huang, Gui Geng, Shengyi Hua, Zhen Huang, Haoyang Zou, Shaoting Zhang, Pengfei Liu, and Xiaofan Zhang. O1 Replication Journey–Part 3: Inference-time Scaling for Medical Reasoning. arXiv preprint arXiv:2501.06458, 2025

  19. [25]

    Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning

    Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, et al. Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning. arXiv preprint arXiv:2410.02884, 2024

  20. [26]

    A survey of monte carlo tree search methods

    Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in...

  21. [27]

    Training Verifiers to Solve Math Word Problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021

  22. [28]

    Measuring mathematical problem solving with the math dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. In Proc. NeurIPS D&B Track, 2021

  23. [29]

    MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline

    Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline. In Findings of Proc. ACL, page 905–924, 2024

  24. [30]

    Direct preference optimization: Your language model is secretly a reward model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Proc. NeurIPS, 2023

  25. [31]

    Proximal policy optimization algorithms

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017

  26. [32]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

  27. [33]

    Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

    Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv preprint arXiv:2502.05171, 2025

  28. [34]

    Training large language models to reason in a continuous latent space

    Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space. In ICLR Workshop on LLM Reason and Plan, 2024

  29. [35]

    Deepseek-v3 technical report

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024

  30. [36]

    Qwen2.5 technical report

    A Yang Qwen, Baosong Yang, B Zhang, B Hui, B Zheng, B Yu, Chengpeng Li, D Liu, F Huang, H Wei, et al. Qwen2.5 technical report. arXiv preprint, 2024

  31. [37]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv e-prints, pages arXiv–2407, 2024

  32. [38]

    Solving math word problems with process-and outcome-based feedback

    Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022

  33. [39]

    Star: Self-taught reasoner bootstrapping reasoning with reasoning

    Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D Goodman. Star: Self-taught reasoner bootstrapping reasoning with reasoning. In Proc. NeurIPS, volume 1126, 2024

  34. [40]

    Let’s verify step by step

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In Proc. ICLR, 2023

  35. [41]

    T \" ulu 3: Pushing frontiers in open language model post-training

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al. T \" ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024

  36. [42]

    Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation

    Youssef Mroueh. Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation. arXiv preprint arXiv:2503.06639, 2025

  37. [43]

    Perception, reason, think, and plan: A survey on large multimodal reasoning models

    Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. Perception, reason, think, and plan: A survey on large multimodal reasoning models. arXiv preprint arXiv:2505.04921, 2025

  38. [44]

    Multimodal chain-of- thought reasoning: A comprehensive survey

    Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. Multimodal chain-of- thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605, 2025

  39. [45]

    Multimodal chain-of-thought reasoning in language models

    Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research, 2024

  40. [46]

    Video-of-thought: Step-by-step video reasoning from perception to cognition

    Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Proc. ICML, 2024

  41. [47]

    Thinking before looking: Improving multimodal llm reasoning via mitigating visual hallucination

    Haojie Zheng, Tianyang Xu, Hanchi Sun, Shu Pu, Ruoxi Chen, and Lichao Sun. Thinking before looking: Improving multimodal llm reasoning via mitigating visual hallucination. arXiv preprint arXiv:2411.12591, 2024. 25 References A Preprint

  42. [48]

    Llava-o1: Let vision language models reason step-by-step

    Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan. Llava-o1: Let vision language models reason step-by-step. arXiv preprint arXiv:2411.10440, 2024

  43. [49]

    Llamav-o1: Rethinking step-by-step visual reasoning in llms

    Omkar Thawakar, Dinura Dissanayake, Ketan More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, et al. Llamav-o1: Rethinking step-by-step visual reasoning in llms. arXiv preprint arXiv:2501.06186, 2025

  44. [50]

    RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? arXiv preprint arXiv:2501.11284, 2025

    Haotian Xu, Xing Wu, Weinong Wang, Zhongzhi Li, Da Zheng, Boyuan Chen, Yi Hu, Shijia Kang, Jiaming Ji, Yingying Zhang, et al. RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? arXiv preprint arXiv:2501.11284, 2025

  45. [51]

    Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search

    Huanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang, Yibo Wang, Shunyu Liu, Yingjie Wang, Yuxin Song, Haocheng Feng, Li Shen, et al. Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search. arXiv preprint arXiv:2412.18319, 2024

  46. [52]

    Improve vision language model chain-of-thought reasoning

    Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, and Yiming Yang. Improve vision language model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198, 2024

  47. [53]

    Insight-v: Exploring long-chain visual reasoning with multimodal large language models

    Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. In Proc. CVPR, pages 9062–9072, 2025

  48. [54]

    Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale

    Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xiang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237, 2024

  49. [55]

    Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification

    Linzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, and Wentao Zhang. Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification. arXiv preprint arXiv:2502.13383, 2025

  50. [56]

    Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking

    Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking. arXiv preprint arXiv:2501.01306, 2025

  51. [57]

    HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning

    Shayan Ali Akbar, Md Mosharaf Hossain, Tess Wood, Si-Chi Chin, Erica M Salinas, Victor Alvarez, and Erwin Cornejo. HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning. In Proc. EMNLP, pages 15020– 15037, 2024

  52. [58]

    CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection

    Ron Eliav, Arie Cattan, Eran Hirsch, Shahaf Bassan, Elias Stengel-Eskin, Mohit Bansal, and Ido Dagan. CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection. arXiv preprint arXiv:2506.05243, 2025

  53. [59]

    Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language- Models

    Zikai Xie. Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language- Models. arXiv preprint arXiv:2408.05093, 2024

  54. [60]

    Grounded chain-of- thought for multimodal large language models

    Qiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang, Baiyang Song, Xiaoshuai Sun, and Rongrong Ji. Grounded chain-of- thought for multimodal large language models. arXiv preprint arXiv:2503.12799, 2025

  55. [61]

    CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation

    Yue Jiang, Jiawei Chen, Dingkang Yang, Mingcheng Li, Shunli Wang, Tong Wu, Ke Li, and Lihua Zhang. CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation. In Proc. ICASSP, 2025

  56. [62]

    MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

    Bowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, and Lei Zhang. MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM. arXiv preprint arXiv:2505.24238, 2025

  57. [63]

    The Hallucination Tax of Reinforcement Finetuning.arXiv preprint arXiv:2505.13988, 2025

    Linxin Song, Taiwei Shi, and Jieyu Zhao. The Hallucination Tax of Reinforcement Finetuning.arXiv preprint arXiv:2505.13988, 2025

  58. [64]

    More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

    Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou, Xin Eric Wang, Yuyin Zhou, and Sheng Liu. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models. arXiv preprint arXiv:2505.21523, 2025

  59. [65]

    Are Reasoning Models More Prone to Hallucination? arXiv preprint arXiv:2505.23646, 2025

    Zijun Yao, Yantao Liu, Yanxu Chen, Jianhui Chen, Junfeng Fang, Lei Hou, Juanzi Li, and Tat-Seng Chua. Are Reasoning Models More Prone to Hallucination? arXiv preprint arXiv:2505.23646, 2025

  60. [66]

    AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, and Samuel J Bell. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions. arXiv preprint arXiv:2506.09038, 2025

  61. [67]

    Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models

    Haolang Lu, Yilian Liu, Jingxin Xu, Guoshun Nan, Yuanlong Yu, Zhican Chen, and Kun Wang. Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models. arXiv preprint arXiv:2505.13143, 2025

  62. [68]

    The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models

    Junyi Li and Hwee Tou Ng. The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models. arXiv preprint arXiv:2505.24630, 2025

  63. [69]

    Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning

    Dang Hoang Anh, Vu Tran, and Le Minh Nguyen. Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning. In JSAI International Symposium on Artificial Intelligence, pages 179–195. Springer, 2025

  64. [70]

    Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

    Zhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang, and Jun Xu. Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective. arXiv preprint arXiv:2505.12886, 2025

  65. [71]

    Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

    Dadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He, Haoran Li, Yumeng Wang, et al. Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models. arXiv preprint arXiv:2506.17114, 2025

  66. [72]

    Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning

    Ruosen Li, Ziming Luo, and Xinya Du. Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning. arXiv preprint arXiv:2410.06304, 2024. 26 References A Preprint

  67. [73]

    Reasoning Models Know When They’re Right: Probing Hidden States for Self-Verification

    Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, and He He. Reasoning Models Know When They’re Right: Probing Hidden States for Self-Verification. arXiv preprint arXiv:2504.05419, 2025

  68. [74]

    Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models

    Changyue Wang, Weihang Su, Qingyao Ai, and Yiqun Liu. Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models. arXiv preprint arXiv:2506.04832, 2025

  69. [75]

    Measuring faithfulness in chain-of-thought reasoning

    Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702, 2023

  70. [76]

    Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Proc. NeurIPS, 2023

  71. [77]

    Measuring faithfulness of chains of thought by unlearning reasoning steps

    Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasović, and Yonatan Belinkov. Measuring faithfulness of chains of thought by unlearning reasoning steps. arXiv preprint arXiv:2502.14829, 2025

  72. [78]

    Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models

    Zidi Xiong, Chen Shan, Zhenting Qi, and Himabindu Lakkaraju. Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models. arXiv preprint arXiv:2505.13774, 2025

  73. [79]

    Chain-of-Thought Unfaithfulness as Disguised Accuracy

    Oliver Bentham, Nathan Stringham, and Ana Marasovic. Chain-of-Thought Unfaithfulness as Disguised Accuracy. Transac- tions on Machine Learning Research, 2024

  74. [80]

    Chain-of- thought reasoning in the wild is not always faithful

    Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy. Chain-of- thought reasoning in the wild is not always faithful. arXiv preprint arXiv:2503.08679, 2025

  75. [83]

    Towards faithful chain-of-thought: Large language models are bridging reasoners

    Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. Towards faithful chain-of-thought: Large language models are bridging reasoners. arXiv preprint arXiv:2405.18915, 2024

  76. [84]

    Faithfulness vs

    Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. Faithfulness vs. plausibility: On the (un) reliability of explanations from large language models. arXiv preprint arXiv:2402.04614, 2024

  77. [85]

    How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc

    Guangsheng Bao, Hongbo Zhang, Cunxiang Wang, Linyi Yang, and Yue Zhang. How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc. COLING, 2024

  78. [86]

    On the difficulty of faithful chain-of-thought reasoning in large language models

    Sree Harsha Tanneru, Dan Ley, Chirag Agarwal, and Himabindu Lakkaraju. On the difficulty of faithful chain-of-thought reasoning in large language models. In ICML Workshop on TiFA, 2024

  79. [87]

    On the impact of fine-tuning on chain-of-thought reasoning

    Elita Lobo, Chirag Agarwal, and Himabindu Lakkaraju. On the impact of fine-tuning on chain-of-thought reasoning. In Proc. NAACL, 2025

  80. [88]

    Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning

    Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. In Findings of Proc. EMNLP, pages 15012–15032, 2024

  81. [89]

    Faithful logical reasoning via symbolic chain-of-thought

    Jundong Xu, Hao Fei, Liangming Pan, Qian Liu, Mong-Li Lee, and Wynne Hsu. Faithful logical reasoning via symbolic chain-of-thought. In Proc. ACL, 2024

  82. [90]

    Question decomposition improves the faithfulness of model-generated reasoning

    Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamil˙e Lukoši¯ut˙e, et al. Question decomposition improves the faithfulness of model-generated reasoning. arXiv preprint arXiv:2307.11768, 2023

  83. [91]

    Faithful chain-of-thought reasoning

    Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. Faithful chain-of-thought reasoning. In Proc. IJCNLP-AACL, 2023

  84. [92]

    Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

    Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning. In Proc. EMNLP, 2023

  85. [93]

    FLARE: Faithful Logic-Aided Reasoning and Exploration

    Erik Arakelyan, Pasquale Minervini, Pat Verga, Patrick Lewis, and Isabelle Augenstein. FLARE: Faithful Logic-Aided Reasoning and Exploration. arXiv preprint arXiv:2410.11900, 2024

  86. [94]

    CoMAT: Chain of mathematically annotated thought improves mathematical reasoning

    Joshua Ong Jun Leang, Aryo Pradipta Gema, and Shay B Cohen. CoMAT: Chain of mathematically annotated thought improves mathematical reasoning. arXiv preprint arXiv:2410.10336, 2024

  87. [95]

    Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering

    Jiawei Wang, Da Cao, Shaofei Lu, Zhanchang Ma, Junbin Xiao, and Tat-Seng Chua. Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering. In Proc. MM, pages 4331–4340, 2024

  88. [96]

    Fact: Teaching mllms with faithful, concise and transferable rationales

    Minghe Gao, Shuang Chen, Liang Pang, Yuan Yao, Jisheng Dang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang, and Tat-Seng Chua. Fact: Teaching mllms with faithful, concise and transferable rationales. In Proc. MM, pages 846–855, 2024

  89. [97]

    Markovian Transformers for Informative Language Modeling

    Scott Viteri, Max Lamparth, Peter Chatain, and Clark Barrett. Markovian Transformers for Informative Language Modeling. arXiv preprint arXiv:2404.18988, 2024. 27 References A Preprint

  90. [98]

    Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts

    Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Limin Han, Jiaojiao Zhao, Beibei Huang, Zhenhong Long, Junting Guo, Meijuan An, Rongjia Du, et al. Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts. arXiv preprint arXiv:2503.16529, 2025

  91. [99]

    Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

    Miguel Romero-Arjona, Pablo Valle, Juan C Alonso, Ana B Sánchez, Miriam Ugarte, Antonia Cazalilla, Vicente Cambrón, José A Parejo, Aitor Arrieta, and Sergio Segura. Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives. arXiv preprint arXiv:2503.10192, 2025

  92. [100]

    The hidden risks of large reasoning models: A safety assessment of r1

    Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Shreedhar Jangam, Jayanth Srinivasa, Gaowen Liu, Dawn Song, and Xin Eric Wang. The hidden risks of large reasoning models: A safety assessment of r1. arXiv preprint arXiv:2502.12659, 2025

  93. [101]

    Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning

    Ang Li, Yichuan Mo, Mingjie Li, Yifei Wang, and Yisen Wang. Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning. arXiv preprint arXiv:2502.09673, 2025

  94. [102]

    Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model

    Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, and Kaiyu Huang. Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model. arXiv preprint arXiv:2505.06538, 2025

  95. [103]

    Evaluating Security Risk in DeepSeek and Other Frontier Reasoning Models

    Paul Kassianik and Amin Karbasi. Evaluating Security Risk in DeepSeek and Other Frontier Reasoning Models. Cisco, https://blogs. cisco. com/security/evaluating-security-risk-in-deepseek-and-other-frontier-reasoningmodels, 2025

  96. [104]

    Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

    Arjun Krishna, Aaditya Rastogi, and Erick Galinkin. Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models. arXiv preprint arXiv:2506.13726, 2025

  97. [105]

    FORTRESS: Frontier Risk Evaluation for National Security and Public Safety

    Christina Q Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Team, and Julian Michael. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety. arXiv preprint arXiv:2506.14922, 2025

  98. [106]

    Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems

    Yihe Fan, Wenqi Zhang, Xudong Pan, and Min Yang. Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems. arXiv preprint arXiv:2505.17815, 2025

  99. [107]

    Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

    Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Wenbo Su, Xiaoyong Zhu, et al. Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models. arXiv preprint arXiv:2505.19690, 2025

  100. [108]

    IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks

    Xiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou, Weichen Zhang, Dongrui Liu, Lu Sheng, and Jing Shao. IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks. arXiv preprint arXiv:2506.16402, 2025

  101. [109]

    SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models

    Junfeng Fang, Yukai Wang, Ruipeng Wang, Zijun Yao, Kun Wang, An Zhang, Xiang Wang, and Tat-Seng Chua. SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models. arXiv preprint arXiv:2504.08813, 2025

  102. [110]

    Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over foundational capabilities

    Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che, Tat-Seng Chua, and Ting Liu. Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over foundational capabilities. arXiv preprint arXiv:...

  103. [111]

    DeepSeek-R1 Thoughtology: Let’s think about LLM Reasoning

    Sara Vera Marjanović, Arkil Patel, Vaibhav Adlakha, Milad Aghajohari, Parishad BehnamGhader, Mehar Bhatia, Aditi Khandelwal, Austin Kraft, Benno Krojer, Xing Han Lù, et al. DeepSeek-R1 Thoughtology: Let’s think about LLM Reasoning. arXiv preprint arXiv:2504.07128, 2025

  104. [112]

    Adversarial Reasoning at Jailbreaking Time

    Mahdi Sabbaghi, Paul Kassianik, George Pappas, Yaron Singer, Amin Karbasi, and Hamed Hassani. Adversarial Reasoning at Jailbreaking Time. In Proc. ICML, 2025

  105. [113]

    Enhancing Adversarial Attacks through Chain of Thought

    Jingbo Su. Enhancing Adversarial Attacks through Chain of Thought. arXiv preprint arXiv:2410.21791, 2024

  106. [114]

    Reasoning-augmented conversation for multi-turn jailbreak attacks on large language models

    Zonghao Ying, Deyue Zhang, Zonglei Jing, Yisong Xiao, Quanchen Zou, Aishan Liu, Siyuan Liang, Xiangzheng Zhang, Xianglong Liu, and Dacheng Tao. Reasoning-augmented conversation for multi-turn jailbreak attacks on large language models. arXiv preprint arXiv:2502.11054, 2025

  107. [115]

    Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models

    Wenhan Chang, Tianqing Zhu, Yu Zhao, Shuangyong Song, Ping Xiong, Wanlei Zhou, and Yongxiang Li. Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models. arXiv preprint arXiv:2505.17519, 2025

  108. [116]

    competency

    Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, and Chitta Baral. When “competency" in reasoning opens the door to vulnerability: Jailbreaking llms via novel complex ciphers. arXiv preprint arXiv:2402.10601, 2024

  109. [117]

    H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking

    Martin Kuo, Jianyi Zhang, Aolin Ding, Qinsi Wang, Louis DiValentin, Yujia Bao, Wei Wei, Hai Li, and Yiran Chen. H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash think...

  110. [118]

    A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos

    Yang Yao, Xuan Tong, Ruofan Wang, Yixu Wang, Lujundong Li, Liang Liu, Yan Teng, and Yingchun Wang. A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos. arXiv preprint arXiv:2502.15806, 2025

  111. [119]

    AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Models

    Jiacheng Liang, Tanqiu Jiang, Yuhui Wang, Rongyi Zhu, Fenglong Ma, and Ting Wang. AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Models. arXiv preprint arXiv:2505.10846, 2025

  112. [120]

    Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers

    Viet-Anh Nguyen, Shiqian Zhao, Gia Dao, Runyi Hu, Yi Xie, and Luu Anh Tuan. Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers. arXiv preprint arXiv:2505.16241, 2025

  113. [121]

    Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

    Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, and Lap-Pui Chau. Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models. arXiv preprint arXiv:2504.05050, 2025

  114. [122]

    RRTL: Red Teaming Reasoning Large Language Models in Tool Learning.arXiv preprint arXiv:2505.17106, 2025

    Yifei Liu, Yu Cui, and Haibin Zhang. RRTL: Red Teaming Reasoning Large Language Models in Tool Learning.arXiv preprint arXiv:2505.17106, 2025. 28 References A Preprint

  115. [123]

    VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models

    Bingrui Sima, Linhua Cong, Wenxuan Wang, and Kun He. VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models. arXiv preprint arXiv:2505.19684, 2025

  116. [124]

    HauntAttack: When Attack Follows Reasoning as a Shadow

    Jingyuan Ma, Rui Li, Zheng Li, Junfeng Liu, Lei Sha, and Zhifang Sui. HauntAttack: When Attack Follows Reasoning as a Shadow. arXiv preprint arXiv:2506.07031, 2025

  117. [125]

    GuardReasoner: Towards Reasoning-based LLM Safeguards

    Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, and Bryan Hooi. GuardReasoner: Towards Reasoning-based LLM Safeguards. arXiv preprint arXiv:2501.18492, 2025

  118. [126]

    X-Guard: Multilingual guard agent for content moderation

    Bibek Upadhayay, Vahid Behzadan, et al. X-Guard: Multilingual guard agent for content moderation. arXiv preprint arXiv:2504.08848, 2025

  119. [127]

    Yahan Yang, Soham Dan, Shuo Li, Dan Roth, and Insup Lee. MR. Guard: Multilingual Reasoning Guardrail using Curriculum Learning. arXiv preprint arXiv:2504.15241, 2025

  120. [128]

    RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards.arXiv preprint arXiv:2506.07736, 2025

    Jingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang, and Tat-Seng Chua. RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards.arXiv preprint arXiv:2506.07736, 2025

  121. [129]

    Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models

    Makesh Narsimhan Sreedhar, Traian Rebedea, and Christopher Parisien. Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models. arXiv preprint arXiv:2505.20087, 2025

  122. [130]

    R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning

    Mintong Kang and Bo Li. R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning. In Proc. ICLR, 2025

  123. [132]

    ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

    Shiyao Cui, Qinglin Zhang, Xuan Ouyang, Renmiao Chen, Zhexin Zhang, Yida Lu, Hongning Wang, Han Qiu, and Minlie Huang. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs. arXiv preprint arXiv:2505.14035, 2025

  124. [133]

    Guardreasoner-vl: Safeguarding vlms via reinforced reasoning

    Yue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang, Xinfeng Li, Kun Wang, Junfeng Fang, et al. Guardreasoner-vl: Safeguarding vlms via reinforced reasoning. arXiv preprint arXiv:2505.11049, 2025

  125. [134]

    GuardAgent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning

    Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, et al. GuardAgent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning. arXiv preprint arXiv:2406.09187, 2024

  126. [135]

    ShieldAgent: Shielding agents via verifiable safety policy reasoning

    Zhaorun Chen, Mintong Kang, and Bo Li. ShieldAgent: Shielding agents via verifiable safety policy reasoning. arXiv preprint arXiv:2503.22738, 2025

  127. [136]

    Unified multimodal chain-of- thought reward model through reinforcement fine-tuning

    Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, and Jiaqi Wang. Unified multimodal chain-of- thought reward model through reinforcement fine-tuning. arXiv preprint arXiv:2505.03318, 2025

  128. [137]

    Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning

    Fengjun Pan, Anh Tuan Luu, and Xiaobao Wu. Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning. arXiv preprint arXiv:2506.08477, 2025

  129. [138]

    Effectively Controlling Reasoning Models through Thinking Intervention

    Tong Wu, Chong Xiang, Jiachen T Wang, and Prateek Mittal. Effectively Controlling Reasoning Models through Thinking Intervention. arXiv preprint arXiv:2503.24370, 2025

  130. [139]

    Adversarial Manipulation of Reasoning Models using Internal Representations

    Kureha Yamaguchi, Benjamin Etheridge, and Andy Arditi. Adversarial Manipulation of Reasoning Models using Internal Representations. In ICML 2025 Workshop on Reliable and Responsible Foundation Models, 2025

  131. [140]

    Trading inference-time compute for adversarial robustness

    Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak, Stephanie Lin, Sam Toyer, Yaodong Yu, Rachel Dias, Eric Wallace, Kai Xiao, Johannes Heidecke, et al. Trading inference-time compute for adversarial robustness. arXiv preprint arXiv:2501.18841, 2025

  132. [141]

    Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance

    Ruizhong Qiu, Gaotang Li, Tianxin Wei, Jingrui He, and Hanghang Tong. Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance. arXiv preprint arXiv:2506.06444, 2025

  133. [142]

    Mixture of insightful experts (mote): The synergy of thought chains and expert mixtures in self-alignment

    Zhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, et al. Mixture of insightful experts (mote): The synergy of thought chains and expert mixtures in self-alignment. arXiv preprint arXiv:2405.00557, 2024

  134. [143]

    Backtracking improves generation safety

    Yiming Zhang, Jianfeng Chi, Hailey Nguyen, Kartikeya Upasani, Daniel M Bikel, Jason Weston, and Eric Michael Smith. Backtracking improves generation safety. In Proc. ICLR, 2025

  135. [144]

    Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning

    Xianglin Yang, Gelei Deng, Jieming Shi, Tianwei Zhang, and Jin Song Dong. Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning. arXiv preprint arXiv:2501.19180, 2025

  136. [145]

    Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking

    Junda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin, and Lei Sha. Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking. arXiv preprint arXiv:2502.12970, 2025

  137. [146]

    Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety

    Yuyou Zhang, Miao Li, William Han, Yihang Yao, Zhepeng Cen, and Ding Zhao. Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety. arXiv preprint arXiv:2503.05021, 2025

  138. [147]

    ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimization

    Kehua Feng, Keyan Ding, Jing Yu, Menghan Li, Yuhao Wang, Tong Xu, Xinda Wang, Qiang Zhang, and Huajun Chen. ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimization. arXiv preprint arXiv:2504.02725, 2025. 29 References A Preprint

  139. [148]

    SaRO: Enhancing LLM Safety through Reasoning-based Alignment

    Yutao Mou, Yuxiao Luo, Shikun Zhang, and Wei Ye. SaRO: Enhancing LLM Safety through Reasoning-based Alignment. arXiv preprint arXiv:2504.09420, 2025

  140. [149]

    Reasoning as an Adaptive Defense for Safety.arXiv preprint arXiv:2507.00971, 2025

    Taeyoun Kim, Fahim Tajwar, Aditi Raghunathan, and Aviral Kumar. Reasoning as an Adaptive Defense for Safety.arXiv preprint arXiv:2507.00971, 2025

  141. [150]

    Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

    Changyue Jiang, Xudong Pan, and Min Yang. Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction. arXiv preprint arXiv:2505.11063, 2025

  142. [151]

    ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Models

    Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, and Min Yang. ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Models. arXiv preprint arXiv:2505.17244, 2025

  143. [152]

    Deliberative alignment: Reasoning enables safer language models

    Melody Y Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, et al. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339, 2024

  144. [153]

    STAR-1: Safer Alignment of Reasoning LLMs with 1K Data

    Zijun Wang, Haoqin Tu, Yuhan Wang, Juncheng Wu, Jieru Mei, Brian R Bartoldson, Bhavya Kailkhura, and Cihang Xie. STAR-1: Safer Alignment of Reasoning LLMs with 1K Data. arXiv preprint arXiv:2504.01903, 2025

  145. [154]

    RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

    Yichi Zhang, Zihao Zeng, Dongbai Li, Yao Huang, Zhijie Deng, and Yinpeng Dong. RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability. arXiv preprint arXiv:2504.10081, 2025

  146. [155]

    SAFEPATH: Preventing Harmful Reasoning in Chain-of- Thought via Early Alignment

    Wonje Jeung, Sangyeon Yoon, Minsuk Kahng, and Albert No. SAFEPATH: Preventing Harmful Reasoning in Chain-of- Thought via Early Alignment. arXiv preprint arXiv:2505.14667, 2025

  147. [156]

    Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

    Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Heli Xu, Tianshu Chu, Peizhao Hu, and Yangqiu Song. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning. arXiv preprint arXiv:2505.14585, 2025

  148. [157]

    Monitoring reasoning models for misbehavior and the risks of promoting obfuscation

    Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. arXiv preprint arXiv:2503.11926, 2025

  149. [158]

    How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

    Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, et al. How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study. arXiv preprint arXiv:2505.15404, 2025

  150. [159]

    Hair: Hardness-aware inverse reinforcement learning with introspective reasoning for llm alignment

    Ruoxi Cheng, Haoxuan Ma, and Weixin Wang. Hair: Hardness-aware inverse reinforcement learning with introspective reasoning for llm alignment. arXiv preprint arXiv:2503.18991, 2025

  151. [160]

    Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

    Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, and Natasha Jaques. Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models. arXiv preprint arXiv:2506.07468, 2025

  152. [161]

    Safety tax: Safety alignment makes your large reasoning models less reasonable

    Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. Safety tax: Safety alignment makes your large reasoning models less reasonable. arXiv preprint arXiv:2503.00555, 2025

  153. [162]

    SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning

    Kaiwen Zhou, Xuandong Zhao, Gaowen Liu, Jayanth Srinivasa, Aosong Feng, Dawn Song, and Xin Eric Wang. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning. arXiv preprint arXiv:2505.16186, 2025

  154. [163]

    SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation

    Naizhu Jin, Zhong Li, Yinggang Guo, Chao Su, Tian Zhang, and Qingkai Zeng. SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation. arXiv preprint arXiv:2412.05829, 2024

  155. [164]

    To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models

    Zihao Zhu, Hongbao Zhang, Ruotong Wang, Ke Xu, Siwei Lyu, and Baoyuan Wu. To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models. arXiv preprint arXiv:2502.12202, 2025

  156. [165]

    Shadowcot: Cognitive hijacking for stealthy reasoning backdoors in llms

    Gejian Zhao, Hanzhou Wu, Xinpeng Zhang, and Athanasios V Vasilakos. Shadowcot: Cognitive hijacking for stealthy reasoning backdoors in llms. arXiv preprint arXiv:2504.05605, 2025

  157. [166]

    Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

    James Chua, Jan Betley, Mia Taylor, and Owain Evans. Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models. arXiv preprint arXiv:2506.13206, 2025

  158. [167]

    Badchain: Backdoor chain-of-thought prompting for large language models

    Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. Badchain: Backdoor chain-of-thought prompting for large language models. In Proc. ICLR, 2024

  159. [168]

    Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models

    Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun. Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models. arXiv preprint arXiv:2408.12798, 2024

  160. [169]

    Darkmind: Latent chain-of-thought backdoor in customized llms

    Zhen Guo and Reza Tourani. Darkmind: Latent chain-of-thought backdoor in customized llms. arXiv preprint arXiv:2501.18617, 2025

  161. [170]

    Process or result? manipulated ending tokens can mislead reasoning llms to ignore the correct reasoning steps

    Yu Cui, Bryan Hooi, Yujun Cai, and Yiwei Wang. Process or result? manipulated ending tokens can mislead reasoning llms to ignore the correct reasoning steps. arXiv preprint arXiv:2503.19326, 2025

  162. [171]

    System prompt poisoning: Persistent attacks on large language models beyond user injection

    Jiawei Guo and Haipeng Cai. System prompt poisoning: Persistent attacks on large language models beyond user injection. arXiv preprint arXiv:2505.06493, 2025

  163. [172]

    Practical Reasoning Interruption Attacks on Reasoning Large Language Models

    Yu Cui and Cong Zuo. Practical Reasoning Interruption Attacks on Reasoning Large Language Models. arXiv preprint arXiv:2505.06643, 2025. 30 References A Preprint

  164. [173]

    Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression

    Yu Cui, Yujun Cai, and Yiwei Wang. Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression. arXiv preprint arXiv:2504.20493, 2025

  165. [174]

    Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems

    Hongru Song, Yu-an Liu, Ruqing Zhang, Jiafeng Guo, and Yixing Fan. Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems. arXiv preprint arXiv:2505.16367, 2025

  166. [175]

    Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection

    Ryan Marinelli, Josef Pichlmeier, and Tamas Bisztray. Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection. arXiv preprint arXiv:2503.21464, 2025

  167. [176]

    GUARD: Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generation

    Naizhu Jin, Zhong Li, Tian Zhang, and Qingkai Zeng. GUARD: Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generation. arXiv preprint arXiv:2505.21425, 2025

  168. [177]

    Assessing Judging Bias in Large Reasoning Models: An Empirical Study

    Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Xuandong Zhao, Wenxuan Zhang, Dawn Song, and Bingsheng He. Assessing Judging Bias in Large Reasoning Models: An Empirical Study. arXiv preprint arXiv:2504.09946, 2025

  169. [178]

    Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption

    Wenxiao Wang, Parsa Hosseini, and Soheil Feizi. Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption. arXiv preprint arXiv:2504.20769, 2025

  170. [179]

    Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems? arXiv preprint arXiv:2504.00509, 2025

    Kai Yan, Yufei Xu, Zhengyin Du, Xuesong Yao, Zheyu Wang, Xiaowen Guo, and Jiecao Chen. Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems? arXiv preprint arXiv:2504.00509, 2025

  171. [180]

    Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector

    Haoyan Yang, Runxue Bao, Cao Xiao, Jun Ma, Parminder Bhatia, Shangqian Gao, and Taha Kass-Hout. Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector. arXiv preprint arXiv:2505.17100, 2025

  172. [181]

    Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models

    Yuqing Wang and Yun Zhao. Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models. arXiv preprint arXiv:2406.11020, 2024

  173. [182]

    A Closer Look at System Prompt Robustness

    Norman Mu, Jonathan Lu, Michael Lavery, and David Wagner. A Closer Look at System Prompt Robustness. arXiv preprint arXiv:2502.12197, 2025

  174. [183]

    A Frustratingly Simple Yet Highly Effective At- tack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.arXiv preprint arXiv:2503.10635, 2025

    Zhaoyi Li, Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui, and Zhiqiang Shen. A Frustratingly Simple Yet Highly Effective At- tack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.arXiv preprint arXiv:2503.10635, 2025

  175. [184]

    Reasoning Models Are More Easily Gaslighted Than You Think

    Bin Zhu, Hailong Yin, Jingjing Chen, and Yu-Gang Jiang. Reasoning Models Are More Easily Gaslighted Than You Think. arXiv preprint arXiv:2506.09677, 2025

  176. [185]

    Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? In Proc

    Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo, Zengmao Wang, and Bo Han. Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? In Proc. NeurIPS, 2024

  177. [186]

    Stepwise Reasoning Disruption Attack of LLMs

    Jingyu Peng, Maolin Wang, Xiangyu Zhao, Kai Zhang, Wanyu Wang, Pengyue Jia, Qidong Liu, Ruocheng Guo, and Qi Liu. Stepwise Reasoning Disruption Attack of LLMs. In Proc. ACL, pages 5040–5058, 2025

  178. [187]

    Polymath: Evaluating mathematical reasoning in multilingual contexts

    Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, et al. Polymath: Evaluating mathematical reasoning in multilingual contexts. arXiv preprint arXiv:2504.18428, 2025

  179. [188]

    Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models

    Meghana Rajeev, Rajkumar Ramamurthy, Prapti Trivedi, Vikas Yadav, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudan, James Zou, and Nazneen Rajani. Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models. arXiv preprint arXiv:2503.01781, 2025

  180. [189]

    Benchmarking reasoning robustness in large language models

    Tong Yu, Yongcheng Jing, Xikun Zhang, Wentao Jiang, Wenjie Wu, Yingjie Wang, Wenbin Hu, Bo Du, and Dacheng Tao. Benchmarking reasoning robustness in large language models. arXiv preprint arXiv:2503.04550, 2025

  181. [190]

    MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations

    Kaixuan Huang, Jiacheng Guo, Zihao Li, Xiang Ji, Jiawei Ge, Wenzhe Li, Yingqing Guo, Tianle Cai, Hui Yuan, Runzhe Wang, et al. MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations. arXiv preprint arXiv:2502.06453, 2025

  182. [191]

    CODECRASH: Stress Testing LLM Reasoning under Structural and Semantic Perturbations

    Man Ho Lam, Chaozheng Wang, Jen-tse Huang, and Michael R Lyu. CODECRASH: Stress Testing LLM Reasoning under Structural and Semantic Perturbations. arXiv preprint arXiv:2504.14119, 2025

  183. [192]

    Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation

    Jaechul Roh, Varun Gandhi, Shivani Anilkumar, and Arin Garg. Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation. arXiv preprint arXiv:2506.06971, 2025

  184. [193]

    Preemptive answer “attacks" on chain-of-thought reasoning

    Rongwu Xu, Zehan Qi, and Wei Xu. Preemptive answer “attacks" on chain-of-thought reasoning. In Findings of Proc. ACL, 2024

  185. [194]

    Large language models are unconscious of unreasonability in math problems

    Jingyuan Ma, Damai Dai, Lei Sha, and Zhifang Sui. Large language models are unconscious of unreasonability in math problems. arXiv preprint arXiv:2403.19346, 2024

  186. [195]

    Dnr bench: Benchmarking over-reasoning in reasoning llms

    Masoud Hashemi, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudhan, Jishnu Sethumadhavan Nair, Aman Tiwari, and Vikas Yadav. Dnr bench: Benchmarking over-reasoning in reasoning llms. arXiv preprint arXiv:2503.15793, 2025

  187. [196]

    Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? In Proc

    Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, et al. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? In Proc. ACL, 2025

  188. [197]

    Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? arXiv preprint arXiv:2504.06514, 2025

    Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou. Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? arXiv preprint arXiv:2504.06514, 2025. 31 References A Preprint

  189. [198]

    Excessive Reasoning Attack on Reasoning LLMs

    Wai Man Si, Mingjie Li, Michael Backes, and Yang Zhang. Excessive Reasoning Attack on Reasoning LLMs. arXiv preprint arXiv:2506.14374, 2025

  190. [199]

    Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

    Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, et al. Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs. arXiv preprint arXiv:2501.18585, 2025

  191. [200]

    Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms

    Jinyan Su, Jennifer Healey, Preslav Nakov, and Claire Cardie. Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms. arXiv preprint arXiv:2505.00127, 2025

  192. [201]

    Internal Bias in Reasoning Models leads to Overthinking

    Renfei Dang, Shujian Huang, and Jiajun Chen. Internal Bias in Reasoning Models leads to Overthinking. arXiv preprint arXiv:2505.16448, 2025

  193. [202]

    Overthink: Slowdown attacks on reasoning llms

    Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, and Eugene Bagdasarian. Overthink: Slowdown attacks on reasoning llms. arXiv preprint arXiv:2502.02542, 2025

  194. [203]

    The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

    Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al. The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks. arXiv preprint arXiv:2502.08235, 2025

  195. [204]

    Output Length Effect on DeepSeek-R1’s Safety in Forced Thinking

    Xuying Li, Zhuo Li, Yuji Kosuga, and Victor Bian. Output Length Effect on DeepSeek-R1’s Safety in Forced Thinking. arXiv preprint arXiv:2503.01923, 2025

  196. [205]

    ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models

    Chung-En Sun, Ge Yan, and Tsui-Wei Weng. ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models. arXiv preprint arXiv:2503.22048, 2025

  197. [206]

    Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

    Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J Wooldridge, Janet B Pierrehumbert, and Furu Wei. Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks. In Proc. ACL, 2025

  198. [207]

    Detection, Classification, and Mitigation of Gender Bias in Large Language Models

    Xiaoqing Cheng, Hongying Zan, Lulu Kong, Jinwang Song, and Min Peng. Detection, Classification, and Mitigation of Gender Bias in Large Language Models. arXiv preprint arXiv:2506.12527, 2025

  199. [208]

    Prompting techniques for reducing social bias in llms through system 1 and system 2 cognitive processes

    Mahammed Kamruzzaman and Gene Louis Kim. Prompting techniques for reducing social bias in llms through system 1 and system 2 cognitive processes. arXiv preprint arXiv:2404.17218, 2024

  200. [209]

    Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

    Saloni Dash, Amélie Reymond, Emma S Spiro, and Aylin Caliskan. Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning. arXiv preprint arXiv:2506.20020, 2025

  201. [210]

    Bias runs deep: Implicit reasoning biases in persona-assigned llms

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: Implicit reasoning biases in persona-assigned llms. In Proc. ICLR, 2024

  202. [211]

    Biasguard: A reasoning-enhanced bias detection tool for large language models

    Zhiting Fan, Ruizhe Chen, and Zuozhu Liu. Biasguard: A reasoning-enhanced bias detection tool for large language models. In Findings of Proc. ACL, 2025

  203. [212]

    Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models

    Riccardo Cantini, Nicola Gabriele, Alessio Orsino, and Domenico Talia. Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models. arXiv preprint arXiv:2507.02799, 2025

  204. [213]

    R-tofu: Unlearning in large reasoning models

    Sangyeon Yoon, Wonje Jeung, and Albert No. R-tofu: Unlearning in large reasoning models. arXiv preprint arXiv:2505.15214, 2025

  205. [214]

    Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills

    Changsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, and Sijia Liu. Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills. arXiv preprint arXiv:2506.12963, 2025

  206. [215]

    Step-by-Step Reasoning Attack: Revealing ’Erased’ Knowledge in Large Language Models

    Yash Sinha, Manit Baser, Murari Mandal, Dinil Mon Divakaran, and Mohan Kankanhalli. Step-by-Step Reasoning Attack: Revealing ’Erased’ Knowledge in Large Language Models. arXiv preprint arXiv:2506.17279, 2025

  207. [216]

    ImF: Implicit Fingerprint for Large Language Models

    Peng Wanli, Xue Yiming, et al. ImF: Implicit Fingerprint for Large Language Models. arXiv preprint arXiv:2503.21805, 2025

  208. [217]

    CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models

    Zhenzhen Ren, GuoBiao Li, Sheng Li, Zhenxing Qian, and Xinpeng Zhang. CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models. arXiv preprint arXiv:2505.16785, 2025

  209. [218]

    Towards copyright protection for knowledge bases of retrieval-augmented language models via ownership verification with reasoning

    Junfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu, Chenxi Liu, Yanshuo Chen, and Heng Huang. Towards copyright protection for knowledge bases of retrieval-augmented language models via ownership verification with reasoning. arXiv preprint arXiv:2502.10440, 2025

  210. [219]

    Antidistillation sampling

    Yash Savani, Asher Trockman, Zhili Feng, Avi Schwarzschild, Alexander Robey, Marc Finzi, and J Zico Kolter. Antidistillation sampling. arXiv preprint arXiv:2504.13146, 2025

  211. [220]

    Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers

    Tommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun, and Seong Joon Oh. Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers. arXiv preprint arXiv:2506.15674, 2025

  212. [221]

    Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models

    Weidi Luo, Tianyu Lu, Qiming Zhang, Xiaogeng Liu, Bin Hu, Yue Zhao, Jieyu Zhao, Song Gao, Patrick McDaniel, Zhen Xiang, et al. Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models. arXiv preprint arXiv:2504.19373, 2025

  213. [222]

    TrustLLM: Trustworthiness in Large Language Models

    Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, et al. TrustLLM: Trustworthiness in Large Language Models. In Proc. ICML, 2024. 32 References A Preprint

  214. [223]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Informatio...

  215. [224]

    A survey of hallucination in large foundation models

    Vipula Rawte, Amit Sheth, and Amitava Das. A survey of hallucination in large foundation models. arXiv preprint arXiv:2309.05922, 2023

  216. [225]

    Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation.arXiv preprint arXiv:2506.17088, 2025

    Jiahao Cheng, Tiancheng Su, Jia Yuan, Guoxiu He, Jiawei Liu, Xinqi Tao, Jingwen Xie, and Huaxia Li. Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation.arXiv preprint arXiv:2506.17088, 2025

  217. [226]

    Scaling llm test-time compute optimally can be more effective than scaling model parameters

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024

  218. [227]

    Self-consistency improves chain of thought reasoning in language models

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In Proc. ICLR, 2023

  219. [228]

    Truthfulqa: Measuring how models mimic human falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Proc. ACL, 2022

  220. [229]

    Halueval: A large-scale hallucination evaluation benchmark for large language models

    Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. Halueval: A large-scale hallucination evaluation benchmark for large language models. In Proc. EMNLP, 2023

  221. [230]

    Evaluating hallucinations in chinese large language models

    Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, et al. Evaluating hallucinations in chinese large language models. arXiv preprint arXiv:2310.03368, 2023

  222. [231]

    Measuring short-form factuality in large language models

    Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368, 2024

  223. [232]

    Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

    Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017

  224. [233]

    Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods

    Wei Li, Wenhao Wu, Moye Chen, Jiachen Liu, Xinyan Xiao, and Hua Wu. Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods. arXiv preprint arXiv:2203.05227, 2022

  225. [234]

    Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? In Proc

    Alon Jacovi and Yoav Goldberg. Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? In Proc. ACL, pages 4198–4205, 2020

  226. [235]

    Dissociation of faithful and unfaithful reasoning in llms

    Evelyn Yee, Alice Li, Chenyu Tang, Yeon Ho Jung, Ramamohan Paturi, and Leon Bergen. Dissociation of faithful and unfaithful reasoning in llms. arXiv preprint arXiv:2405.15092, 2024

  227. [236]

    Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language? In Findings of Proc

    Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language? In Findings of Proc. EMNLP, pages 4351–4367, 2020

  228. [237]

    Negative preference optimization: From catastrophic collapse to effective unlearning

    Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei. Negative preference optimization: From catastrophic collapse to effective unlearning. In Proc. COLM, 2024

  229. [238]

    Causality

    Judea Pearl. Causality. Cambridge university press, 2009

  230. [239]

    Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

    Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Findings of Proc. ACL, p...

  231. [240]

    Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

    Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. In Proc. ICML, 2024

  232. [241]

    A strongreject for empty jailbreaks

    Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al. A strongreject for empty jailbreaks. In Proc. NeurIPS D&B Track, 2024

  233. [242]

    Air-bench 2024: A safety benchmark based on risk categories from regulations and policies

    Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, et al. Air-bench 2024: A safety benchmark based on risk categories from regulations and policies. arXiv preprint arXiv:2407.17436, 2024

  234. [243]

    WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

    Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs. In Proc. NeurIPS D&B Track, 2024

  235. [244]

    Universal and transferable adversarial attacks on aligned language models

    Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023

  236. [245]

    Jailbreaking black box large language models in twenty queries

    Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In Proc. SaTML, 2025

  237. [246]

    Tree of attacks: Jailbreaking black-box llms automatically

    Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. In Proc. NeurIPS, 2024

  238. [247]

    Gemini 2.0 Flash Thinking, 2025

    Google DeepMind. Gemini 2.0 Flash Thinking, 2025. 33 References A Preprint

  239. [248]

    Kimi k1.5: Scaling reinforcement learning with llms

    Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025

  240. [249]

    Sky-T1: Train your own O1 preview model within $450

    NovaSky Team. Sky-T1: Train your own O1 preview model within $450. https://novasky-ai.github.io/posts/sky-t1, 2025. Accessed: 2025-01-09

  241. [250]

    QwQ-32B: Embracing the Power of Reinforcement Learning, March 2025

    Qwen Team. QwQ-32B: Embracing the Power of Reinforcement Learning, March 2025

  242. [251]

    Skywork-o1 Open Series

    Skywork o1 Team. Skywork-o1 Open Series. https://huggingface.co/Skywork, November 2024

  243. [252]

    Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

    Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models. In Proc. NeurIPS, 2024

  244. [253]

    Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models

    Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, et al. Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models. arXiv...

  245. [254]

    Chisafety- bench: A chinese hierarchical safety benchmark for large language models

    Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An, Bikun Yang, KaiKai Zhao, Kai Wang, and Shiguo Lian. Chisafety- bench: A chinese hierarchical safety benchmark for large language models. arXiv preprint arXiv:2406.10311, 2024

  246. [255]

    Jailbroken: How does llm safety training fail? In Proc

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? In Proc. NeurIPS, 2023

  247. [256]

    R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

    Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, et al. R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization. arXiv preprint arXiv:2503.10615, 2025

  248. [257]

    Skywork r1v: Pioneering multimodal reasoning with chain-of-thought

    Yi Peng, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, et al. Skywork r1v: Pioneering multimodal reasoning with chain-of-thought. arXiv preprint arXiv:2504.05599, 2025

  249. [258]

    Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl

    Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, and Xu Yang. Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl. arXiv preprint arXiv:2503.07536, 2025

  250. [259]

    Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation

    Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation. In Proc. ECCV, pages 388–404, 2024

  251. [260]

    Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models? arXiv preprint arXiv:2504.10000, 2025

    Yanbo Wang, Jiyang Guan, Jian Liang, and Ran He. Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models? arXiv preprint arXiv:2504.10000, 2025

  252. [261]

    More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models. arXiv preprint arXiv:2302.12173, 2023

  253. [262]

    Defending chatgpt against jailbreak attack via self-reminders

    Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. Defending chatgpt against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12):1486–1496, 2023

  254. [263]

    Defending large language models against jailbreaking attacks through goal prioritization

    Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, and Minlie Huang. Defending large language models against jailbreaking attacks through goal prioritization. In Proc. ACL, 2024

  255. [264]

    Jailbreak and guard aligned language models with only few in-context demonstrations

    Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387, 2023

  256. [265]

    Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks

    Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho. Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks. In Findings of Proc. ACL, 2025

  257. [266]

    Root defence strategies: Ensuring safety of llm at the decoding level

    Xinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang, and Yu Tian. Root defence strategies: Ensuring safety of llm at the decoding level. arXiv preprint arXiv:2410.06809, 2024

  258. [267]

    Safedecoding: Defending against jailbreak attacks via safety-aware decoding

    Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. Safedecoding: Defending against jailbreak attacks via safety-aware decoding. arXiv preprint arXiv:2402.08983, 2024

  259. [268]

    Safeinfer: Context adaptive decoding time safety alignment for large language models

    Somnath Banerjee, Sayan Layek, Soham Tripathy, Shanu Kumar, Animesh Mukherjee, and Rima Hazra. Safeinfer: Context adaptive decoding time safety alignment for large language models. In Proc. AAAI, pages 27188–27196, 2025

  260. [269]

    Safeguarding large language models: A survey

    Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, et al. Safeguarding large language models: A survey. arXiv preprint arXiv:2406.02622, 2024

  261. [270]

    Llama guard: Llm-based input-output safeguard for human-ai conversations

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023

  262. [271]

    Llama guard 3 vision: Safeguarding human-ai image understanding conversations

    Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti. Llama guard 3 vision: Safeguarding human-ai image understanding conversations. arXiv preprint arXiv:2411.10414, 2024

  263. [272]

    Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

    llama Team. Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. https://ai.meta.com/blog/llama- 3-2-connect-2024-vision-edge-mobile-devices/, 2024. Accessed: 2024-09-25. 34 References A Preprint

  264. [273]

    A comprehensive survey of LLM alignment techniques: RLHF, RLAIF, PPO, DPO and more

    Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al. A comprehensive survey of LLM alignment techniques: RLHF, RLAIF, PPO, DPO and more. arXiv preprint arXiv:2407.16216, 2024

  265. [274]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In Proc. NeurIPS, 2022

  266. [275]

    Training a helpful and harmless assistant with reinforcement learning from human feedback

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022

  267. [276]

    Rlaif vs

    Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, et al. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. In Proc. ICML, 2024

  268. [277]

    Beavertails: Towards improved safety alignment of llm via a human-preference dataset

    Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. In Proc. NeurIPS, 2024

  269. [278]

    PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

    Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference. In Proc. ACL, 2025

  270. [279]

    Safety fine-tuning at (almost) no cost: A baseline for vision large language models

    Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. In Proc. ICML, 2024

  271. [280]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. In Proc. ICLR, 2022

  272. [281]

    Jailbreakbench: An open robustness benchmark for jailbreaking large language models

    Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Proc. ...

  273. [282]

    STAIR: Improving Safety Alignment with Introspective Reasoning

    Yichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia, Zhengwei Fang, Xiao Yang, Ranjie Duan, Dong Yan, Yinpeng Dong, and Jun Zhu. STAIR: Improving Safety Alignment with Introspective Reasoning. arXiv preprint arXiv:2502.02384, 2025

  274. [283]

    Salad-bench: A hierarchical and comprehensive safety benchmark for large language models

    Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. In Findings of Proc. ACL, 2024

  275. [284]

    Orca: Progressive learning from complex explanation traces of gpt-4

    Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707, 2023

  276. [285]

    Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation

    Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. In Findings of Proc. EMNLP, 2023

  277. [286]

    Sorry-bench: Systematically evaluating large language model safety refusal behaviors

    Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al. Sorry-bench: Systematically evaluating large language model safety refusal behaviors. In Proc. ICLR, 2025

  278. [287]

    Xstest: A test suite for identifying exaggerated safety behaviours in large language models

    Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. In Proc. NAACL, 2024

  279. [288]

    JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

    Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks. In Proc. COLM, 2024

  280. [289]

    Self-refine: Iterative refinement with self-feedback

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. In Proc. NeurIPS, 2023

  281. [290]

    Simplesafetytests: a test suite for identifying critical safety risks in large language models

    Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk, Rebecca Qian, Anand Kannappan, Scott A Hale, and Paul Röttger. Simplesafetytests: a test suite for identifying critical safety risks in large language models. arXiv preprint arXiv:2311.08370, 2023

  282. [291]

    Tdc 2023 (llm edition): The trojan detection challenge

    Mazeika Mantas, Zou Andy, Mu Norman, Phan Long, Wang Zifan, Yu Chunru, Khoja Adam, Jiang Fengqing, O’Gara Aidan, Sakhaee Ellie, Xiang Zhen, Rajabi Arezoo, Hendrycks Dan, Poovendran Radha, Li Bo, and Forsyth David. Tdc 2023 (llm edition): The trojan detection challenge. In Proc...

  283. [292]

    ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming

    Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li. ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming. arXiv preprint arXiv:2404.08676, 2024

  284. [293]

    Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

    Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, et al. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In Proc. CVPR, pages 13807–13816, 2024

  285. [294]

    Aligning large multimodal models with factually augmented rlhf

    Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. In Findings of Proc. ACL, 2024

  286. [295]

    VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment

    Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, and Qi Liu. VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment. In Proc. EMNLP, 2024. 35 References A Preprint

  287. [296]

    Safe RLHF-V: Safe Reinforcement Learning from Human Feedback in Multimodal Large Language Models

    Jiaming Ji, Xinyu Chen, Rui Pan, Han Zhu, Conghui Zhang, Jiahao Li, Donghai Hong, Boyuan Chen, Jiayi Zhou, Kaile Wang, et al. Safe RLHF-V: Safe Reinforcement Learning from Human Feedback in Multimodal Large Language Models. arXiv preprint arXiv:2503.17682, 2025

  288. [297]

    Mm-rlhf: The next step forward in multimodal llm alignment

    Yi-Fan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu, Peiyan Li, Jianshu Zeng, Wulin Xie, Yang Shi, Huanyu Zhang, Junkang Wu, et al. Mm-rlhf: The next step forward in multimodal llm alignment. In Proc. ICML, 2025

  289. [298]

    Explaining and harnessing adversarial examples

    Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014

  290. [299]

    Mitigating the alignment tax of rlhf

    Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, et al. Mitigating the alignment tax of rlhf. In Proc. EMNLP, 2024

  291. [300]

    Findings of the 2014 workshop on statistical machine translation

    Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the ninth workshop on statis...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.