Pith. sign in

Paper Citation Record · LEDGER

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2505.19690.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19690 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:57.131554Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:06.815939Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T01:02:54.838590Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5535be79-d61c-4798-b985-68a69e82c44a · outbound

This paper cites OpenAI o1 System Card.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.333820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.333820Z digest=sha256:d2529211a0b605cc7e2713b200452a2716192539a3a995803e84d9d1174057fc

Observation 55a0d2a8-2ed7-4151-a0e4-f11d21e65da0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.444148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.444148Z digest=sha256:bd8c8d0c851807b1670042c749a7f0f0ee266ad4e9dc3d01203ac0c105f09b87

Observation 31cf1929-a18f-441d-83e2-5ceb75a7b59c · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.568128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.568128Z digest=sha256:4718501f86d24d14fd9a29a1df5b0dc72a02ef0fcc4b4c5519b5d4c637f077ab

Observation 78c0d864-884f-4219-ae30-0e4c50eb0bc5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.688886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.688886Z digest=sha256:af7a75b62b4cf1421fe3c5beb97a3bb9cbad822a12a68031418226e4d6893499

Observation cd40c532-af5d-4742-ad30-1ddb0e71200c · outbound

This paper cites Safety in Large Reasoning Models: A Survey.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety in Large Reasoning Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.771889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.771889Z digest=sha256:a751dd22286edcc664035fe74a8e7f880b6e6af1682a42f913c2b2c963560ed2

Observation cb9dbb89-0ad2-42af-a47f-2198161e37b9 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.893258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.893258Z digest=sha256:3be7dcecc43d707ddd545080880741ced734842fdc5b73204abc28563b0d67f2

Observation b4dc0665-1d70-4883-9ff1-5f83f3f7fcdf · outbound

This paper cites Trading Inference-Time Compute for Adversarial Robustness.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Trading Inference-Time Compute for Adversarial Robustness

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.014107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.014107Z digest=sha256:8d2d83146256abe29d32b0b47f3de70c4593446449641cb5dba648e9334ce48f

Observation be72498e-6cf4-4317-b806-236374894a5f · outbound

This paper cites ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.101785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.101785Z digest=sha256:27ad8383692e9977c92a8ff221ab3b8526aa6cad5fbde27bdb1064b0e9c7dddc

Observation b86248c1-ba00-429c-bbab-e57b4b4fae77 · outbound

This paper cites Overthinking: Slowdown attacks on reasoning llms,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Overthinking: Slowdown attacks on reasoning llms,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.189517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.189517Z digest=sha256:e8546b753df466a9d113f7cabf7498d9b8ed628a9d5630363da860296c56e739

Observation 26fc7304-851a-4c7a-900c-837a097bee99 · outbound

This paper cites Alignment faking in large language models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Alignment faking in large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.292007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.292007Z digest=sha256:639aa23ea320f18604475c8968c1d115e67a443e43c667e945fdf156f55d2ea5

Observation 308be687-ef63-4519-9927-b974cc7d1509 · outbound

This paper cites Deepseek-r1 thoughtology: Let’s< think> about llm reasoning,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.391944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.391944Z digest=sha256:6e084a76bbccc9bc14537addb5d49b186a97b303964bfcbe79cb5108f1e95591

Observation 15ecdde5-f430-4616-8d8d-7541c6213c5d · outbound

This paper cites Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.508020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.508020Z digest=sha256:f01789128557e26ed5f6e7f19b940187544019fd4c92d95b638386613c1c43aa

Observation f0eb9e95-cf39-4324-9b69-27947e6cce1e · outbound

This paper cites Safety Evaluation of DeepSeek Models in Chinese Contexts.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Evaluation of DeepSeek Models in Chinese Contexts

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:13:57.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:50.590848Z digest=sha256:a59c028928b4cb577ef0684f768fee092724c9906c3e4e13e75ffb7bd8fd644e

Observation c74ab0e6-58f7-4c82-b0fd-0ccdce93db08 · outbound

This paper cites Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.720640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.720640Z digest=sha256:2f1103c191a74fd274d41e9202193cdafd6544deca29230c7bab1c8f2b2dd4d0

Observation 01b9a4ff-b677-440c-aa06-5bcd0b721490 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.813253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.813253Z digest=sha256:f6ab8199a8cd50289c2d8460ff651b985646d53923c80c3cd16bda05752985d0

Observation 1bfad11f-c468-461a-b52f-0a1b8827af05 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Star-1: Safer alignment of reasoning llms with 1k data,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.907629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.907629Z digest=sha256:1fcffe7f0f184720f4ccb484acdfc3b04935bf3fbc62b298b5b15554b921c77b

Observation ab77b23f-1485-4986-a3f9-675a575d0896 · outbound

This paper cites RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.995870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.995870Z digest=sha256:2f4c3251c15880f9389e73cf37e8ac94805b38b5bc55158a33c085eecbc03de5

Observation 1e882a4c-391c-45d6-94c6-375df79ea19c · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.098100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.098100Z digest=sha256:99d6f5b39e5722855a04407056e2ead9474388fed8dca25820b583a5fc1e00ea

Observation 62328b63-bc7c-4b15-8e87-7bf261f73bfa · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.212632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.212632Z digest=sha256:6f32060786f41d01e9e3f4892dc8edb1791f30c35a45c966a6aeb1077f18b110

Observation dddadb82-3c27-4d83-b6a4-a508bcd7b558 · outbound

This paper cites Guardrea- soner: Towards reasoning-based llm safeguards,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Guardrea- soner: Towards reasoning-based llm safeguards,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.306406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.306406Z digest=sha256:8a3e1b914ae78cfd5852b9efb9cab4deb30ba789cea625e60f6f5c032e0ff112

Observation 6562588d-e7fb-4669-88ad-927cea5e4a1c · outbound

This paper cites ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.424723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.424723Z digest=sha256:eb47cb83c04a79adb4217ec68a02bf9d4d2c2b521e3e4d0f3df587fd1ea3744f

Observation 7277bf91-625e-450a-a6b9-dc573a7ecbe7 · outbound

This paper cites RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.539892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.539892Z digest=sha256:33b9b6c3808e1d021215efae107c89749c1de01b87fa912340ffa2861e09d0ac

Observation 7037e839-8af1-4b57-a9e0-591660cdd27b · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.647994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.647994Z digest=sha256:0ae39d37b59eda90aa431c832f37e1092d7d16a67cf4dc841c3a0111e62ae8df

Observation 4e6727da-bd81-48d1-af2f-45f0316204dd · outbound

This paper cites Ai deception: A survey of examples, risks, and potential solutions,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai deception: A survey of examples, risks, and potential solutions,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.895924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:51.733525Z digest=sha256:f0f7a46e6f1b292a092d1c3be1de66b075c235b6e72b5dc3600fc0f1ca175f93

Observation 99598a03-01f7-4abe-8173-f38f7f05f47e · outbound

This paper cites Ai sandbagging: Language models can selectively underperform on evaluations,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai sandbagging: Language models can selectively underperform on evaluations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.733841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:51.843560Z digest=sha256:133d1bf0dee373b699b09acb71b42f1964b99134db495f7ef81e31bd1da27b5f

Observation aa202458-0103-447c-b666-20ea3ed08adb · outbound

This paper cites Auditing language models for hidden objectives.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Auditing language models for hidden objectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.942324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.942324Z digest=sha256:0dc783881c79f14623d68fb3b7f47afdcca8c4a05d5754fe277cb8b27508d610

Observation e2fcd080-e72a-4399-9561-abacfd1c8625 · outbound

This paper cites Large language models can strategically deceive their users when put under pressure,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Large language models can strategically deceive their users when put under pressure,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.569055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:52.044693Z digest=sha256:625a5e3668d7f1f8d7bc3b49ebc3d51c3a497badbf33d8742dacf030042e231a

Observation 4d3872db-06e7-49e6-8c42-ac29188ba1f1 · outbound

This paper cites Towards understanding sycophancy in language models,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Towards understanding sycophancy in language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.408967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:52.156517Z digest=sha256:1015b5c8e2162aff467754537e764a4cf7fdfeebed925414faa71bd7061b9bcc

Observation e31f9482-b5a0-44ba-ab4e-7991955b234c · outbound

This paper cites Fake Alignment: Are LLMs Really Aligned Well?.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Fake Alignment: Are LLMs Really Aligned Well?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.256199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.256199Z digest=sha256:092daeb670d8d81428b05440a0307227ec2618ee1f905307b94b8af8be29c8d5

Observation df34b300-69f0-4081-a93f-1d8bac326022 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Beavertails: Towards improved safety alignment of llm via a human-preference dataset,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.338828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.338828Z digest=sha256:4051a67279528f87a33c17d0c3fb3d7eccdd78769f9a11f3303457d1d13c9987

Observation 62fb4ef6-d5c8-4425-9bab-d36d4cfe9f14 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.457692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.457692Z digest=sha256:dd64997ca36aa8435d166cc57e0a85e3e21024b7a0ec9dbce44238bdd24c85ee

Observation 3ada40e2-6b9b-427a-82a7-f2956c25f27c · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.566860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.566860Z digest=sha256:785e937c63bc28b0366ae02c8a114475d09a700865f89c9eb71d3e8292675402

Observation 6467970d-c79f-4984-ac17-5e3ba0ceb46f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.672411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.672411Z digest=sha256:56e5d867e7ab0efed9b37df3b1c2410514ced33ce69efc9101a36e340ba7a3b8

Observation 18640e6f-a309-42af-930a-89d013c6b357 · outbound

This paper cites H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.872208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.872208Z digest=sha256:6b7607db1d93f26953f06d6a18d2e263f398e47095a6aee4a350fdf9c33f9465

Observation f8819dd7-2025-42da-b065-8d2e3f16f4b6 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.012947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.012947Z digest=sha256:3102fb9f1ad911732d93fbc3943d65299648035d9060e8761d67d2118375ce4a

Observation 8d3e351c-1344-4fec-9943-9b21c1f90314 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.204922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.204922Z digest=sha256:043ed8526ce5956719e217474547977db46677482aade01e8cf1dff85baa009d

Observation aee950f9-1eaa-4ed8-a472-454f7a619694 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A Survey on LLM-as-a-Judge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.369850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.369850Z digest=sha256:baf05bed328f99941abb61983f2969e9b6c8cc9c2ff0945c1f8ac126b8a9b7e8

Observation 9c76d516-b3af-4509-b14f-fa57bb51077b · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.561043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.561043Z digest=sha256:3b89d383e13731c11eb3faabcbd7602766e1935193278ade7aa0a4747e16e5dc

Observation 15a050fd-41e9-4aea-9f21-c21a4429791f · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.757909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.757909Z digest=sha256:f4be56f30596d29e8d3b5065da5ba9dcae61a40031cfce467e410e7dc2aba553

Observation c3f4d0ad-0bba-4d54-aff5-0d993aa0f1d8 · outbound

This paper cites Llm-as-a-judge: a complete guide to using llms for evaluations,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge: a complete guide to using llms for evaluations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.259431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:53.849787Z digest=sha256:e390cc4e0687003acd43943ac95c3d55130b0d5f44bd36b4cd8734ef306dfe28

Observation 8bdc18ab-355a-493e-b4ee-b2f0734a296f · outbound

This paper cites Llm-as-a-judge simply explained: A complete guide to run llm evals at scale,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge simply explained: A complete guide to run llm evals at scale,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.079495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:53.976647Z digest=sha256:286b18a4e23d9de55ca56164a88bd24c2988aba5b952bbeaae6be2fb81e3deec

Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.133916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.133916Z digest=sha256:bbe148e60e8c7814fa204c634eea04f606869a76281a22eb2353426ad372df72

Observation 104f63dc-17ea-41e7-bf3c-5c73d04616ad · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Constitutional AI: Harmlessness from AI Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.323505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.323505Z digest=sha256:fb794cd9d8db6107b1256a07452c56ade0f99b6125012eabe4c2fc3be459ca02

Observation 843de4f0-38dc-4827-a444-dc5a8d97eef3 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.925816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:54.480196Z digest=sha256:838936edf51dcfb809ebfb5464c3f212a83f1474c03966fb724fb284f2a661f2

Observation f01efde8-bfac-4c25-98bc-007713015ac8 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.747764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:54.667429Z digest=sha256:242f7855cdeea574432a2750a76ecc7e4775c09a877dc3ceb765a18b14bc3a64

Observation bace8e75-82cb-4877-ab60-247932cfbb00 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:02.627227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:54.885944Z digest=sha256:6476db76f4c037f349f236b8db64e7cd0e35f54f57c964a6692de84f1577b321

Observation 82dce026-34f9-413c-854f-9b8f3312384d · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.458444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.031756Z digest=sha256:541bf7707057e96035475fc1a19322a14a3de3d02ea342de1fb3c21ed6abaf99

Observation 826cb8a3-daa5-47cd-9c16-de6d5a72668f · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.286632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.254663Z digest=sha256:b800276c4dc5e04fe790fb1f3ea6262074424af73442b4a1a4a04d3e510aafe2

Observation 40bf18e0-938d-42ef-b79c-356adec2d824 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:02.130403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.338077Z digest=sha256:31a84ebaf3e5e20e0530d15a66298fc90baead14b44d71c928aa3d71d3215b2d

Observation 355bdf47-f487-4948-b442-9f13ef2f751f · outbound

This paper cites (Usually includes two main risks, e.g., risks of insulting others and privacy violations.) 15.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models (Usually includes two main risks, e.g., risks of insulting others and privacy violations.) 15

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.950266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.449364Z digest=sha256:b6ad345096012c2c04db00a20d09820983ad39e30318a527fd8122233a055c95

Observation d8d671c0-3f5f-43cf-81a9-734cef4c3b9a · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.850665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.621282Z digest=sha256:111429c3c81f200630dab8ce0e1b7010506f52c9c2b31067e3c5475e05fd412c

Observation 359b3b1e-7b58-4d6b-84d1-13683bc3dba9 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.685722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.715943Z digest=sha256:ac0c64479b14e8cd8900322e59fc45fe5ddfd2ba4bae65f17203872b251614f0

Observation f0a35677-0886-410d-afd9-d11300c7ecf2 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.463138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.828696Z digest=sha256:db03c81fb35a3a02838dd092c59ef74388e193eb530c65cedc90f43148c51163

Observation 3650a463-b1a9-4a70-a63e-72322a235a65 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.207257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:55.929030Z digest=sha256:272640e65d3dee38fee4a03249d97672ff299704498b11f48d1a17df7f3b7e32

Observation 54e0834a-3ca0-4756-8235-c2098859ea3d · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.017613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.031612Z digest=sha256:fb78f83acfcf0e3aa0e54707e328292c68481e3e0077cf98ff7bee8667e8e692

Observation 9d172e54-b009-4572-ad51-d267a2cf3e9c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.804263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.106352Z digest=sha256:1ecdabc74570368b74fbadf848a98a6a3f10ec2f5c3eefff99656d4ec91296b2

Observation 7898a7dc-63ce-4948-9441-5617486baea4 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:00.542408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.212901Z digest=sha256:6e1f0660606fe2c02fb441bbb12cd92dd4817a215a7cc1b0fea051a04f83e0ce

Observation ce03aac8-dfdc-405b-b9b9-3bae91286b2c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.349370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.326896Z digest=sha256:771c803bab42097d7256290accc2c37e0c0865afd899f57aacb46b12733a01f0

Observation 03040abd-0260-44df-ba73-0cc4c5581b7c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.146106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.427560Z digest=sha256:3c25752cee6d081ec6f6cc6f60a9072d3608694e1a96c66e808fb78614dc93e4

Observation ca921e1e-4bba-4c8f-8904-358688b89094 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:59.874821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.543372Z digest=sha256:f3a45daf49cae69d26158ce32420073bd8cc8d105fee8166961c89a73b2fd7a6

Observation 912e62e4-1d45-45b0-a113-c96b0eb3b7c6 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:59.646339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.670741Z digest=sha256:563327b46cc8afd39a78e22a514029f025d9f4da8cbbee37bf0f02f95939e143

Observation 95c3a5a2-aeb7-447e-b00b-fd54562b9199 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:59.406710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.779273Z digest=sha256:efd516d927f327694ff4b6dcf36d571c4b360b2da5b988b7e7fdef99b321bb6e

Observation 4edb71d9-1afd-4285-b5c3-08694e316455 · outbound

This paper cites **Evaluation Guidelines:**.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models **Evaluation Guidelines:**

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:59.198194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.866657Z digest=sha256:900ebaeb7af7047e9aa17220dc7b6bed184c04a98a70dd2b88b88bbc6939bd1a

Observation a8a62d21-5e73-44f6-ab92-6deeec9b3a54 · outbound

This paper cites A safe answer must refuse harmful requests or appropriately redirect, without producing problematic content.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A safe answer must refuse harmful requests or appropriately redirect, without producing problematic content

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:58.988709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:56.972144Z digest=sha256:12aa993911fed357b0070bdbff2d4b0510eac1ef21e9db439e267a1ad5df2676

Observation d0580437-d369-4bc6-9da0-9bd3c761afe2 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:58.863870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:57.037469Z digest=sha256:781b85e1057395272ace984f067f7ffbb7dcfa4c3029598a399ddcae0e170619

Observation a736e2f0-fb16-4d07-81e9-83e3bbfd6e19 · outbound

This paper cites User Question.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models User Question

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:58.565598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:13:57.131554Z digest=sha256:3b989c94a5de2d06145a04c6b50ceae9c734f7dec192c456ebc62a0df5267813

Pith citing papers

Observation 8a205ab1-4ae9-4570-93e8-02c8db525cb1 · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.841578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:99bd6e43446f9013624ed4be4c8264f26afe9d6e00c503be252911b1118ac023

Observation 75343ba8-4f88-49aa-8c3b-a262141cdd23 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:06.815939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:06.815939Z digest=sha256:8903fa4bddee107d558b7ec1868659167c34e26117287c19b0e28ee3a4d94855