Pith. sign in

Paper Citation Record · LEDGER

Safety in Large Reasoning Models: A Survey

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 19 inbound Pith citation observations for arXiv:2504.17704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17704 v3

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:36:48.122448Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:10.981064Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 634f42b0-c071-4ebb-8437-de28733da301 · outbound

This paper cites ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time.

Safety in Large Reasoning Models: A Survey ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.083772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.083772Z digest=sha256:798bc547fd04a026785cb01633a01af2cd2b157d3b290bfe5d6706cfca67d5ca

Observation 98b5ea52-b3df-45db-a343-775bc98530d2 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Safety in Large Reasoning Models: A Survey Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.087814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.087814Z digest=sha256:ff6293837da53d622246510688bdac887b54773d2c9736bd9535486f1dae54f4

Observation e930902f-09d0-40f1-a902-9f1c37ffcb69 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

Safety in Large Reasoning Models: A Survey Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.091820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.091820Z digest=sha256:c5925613f0451dee1eee20e2f961cd5dbc52d6b16dab9ab02c80e5b878a2b869

Observation b17e3b7c-f358-475f-a31b-258e57cd2f06 · outbound

This paper cites Mitigating the Alignment Tax of RLHF.

Safety in Large Reasoning Models: A Survey Mitigating the Alignment Tax of RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.095899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.095899Z digest=sha256:86eaeb5982e1f5114f9c39bc2df2124ae8f8f375eff6817c443dc6bc5e642ccf

Observation bac55384-0b22-4612-b14f-b58f6bb13836 · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Safety in Large Reasoning Models: A Survey SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.099786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.099786Z digest=sha256:39ba6b660f0513a0532a7132feb6b8ec10ce5f79dbb7cafd750f821608feeee4

Observation 0413f2a8-2c4b-409b-8b19-0e4dcf6c27b1 · outbound

This paper cites Advances in Neu- ral Information Processing Systems, 36.

Safety in Large Reasoning Models: A Survey Advances in Neu- ral Information Processing Systems, 36

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.111418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.111418Z digest=sha256:da3d777724c9bfe3db53aed52c93aec157f25bb94ed236d1048f34a224cfb4a5

Observation 9d92f40b-a388-4615-be45-ea29d6fe89bb · outbound

This paper cites Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack.

Safety in Large Reasoning Models: A Survey Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.115014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.115014Z digest=sha256:07020c4c3a4018cbc57143618658cae0e2c955a026025a58560740c7687edbac

Observation 09b11d51-9067-4f7a-bc4d-b70bdb1cbb10 · outbound

This paper cites Recursively Summarizing Books with Human Feedback.

Safety in Large Reasoning Models: A Survey Recursively Summarizing Books with Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.118794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.118794Z digest=sha256:5513d0f5f0d46c961a7b3417824229e06a2ef7d68e7fe4a226188b5bb6d3a635

Observation e339bf25-ac8c-4980-b932-ddadb4c3bd1c · outbound

This paper cites BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models.

Safety in Large Reasoning Models: A Survey BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.122448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.122448Z digest=sha256:5ab264461eda2a5899d3cbf501f617e31e6fcbe26c7a3a017416e176b514a98d

Observation 7da4ea5f-cb66-48b6-9384-98df46a20e6d · outbound

This paper cites Emerging Cyber Attack Risks of Medical AI Agents.

Safety in Large Reasoning Models: A Survey Emerging Cyber Attack Risks of Medical AI Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.107330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.107330Z digest=sha256:87751e007bca69987ab29ae0814ee7092310d76ab219b0686297f6aab0380084

Observation 47af50ac-5a02-4a84-9f50-f7ffb2feb7d6 · outbound

This paper cites Advances in neural in- formation processing systems, 35.

Safety in Large Reasoning Models: A Survey Advances in neural in- formation processing systems, 35

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:36:48.325642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T10:36:48.103733Z digest=sha256:6ac49622e5187b178f5b1e7c8bd25ea359980139e320bb7383f006e7cd13f5f4

Observation d852b77c-8f0c-46d4-9ee6-85a30154f37a · outbound

This paper cites Black-Box Prompt Optimization: Aligning Large Language Models without Model Training.

Safety in Large Reasoning Models: A Survey Black-Box Prompt Optimization: Aligning Large Language Models without Model Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.075165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.075165Z digest=sha256:8de12cff28ea666e67fcf01571e9ba835a3a114a64930912ea7f0f2d62388baa

Observation 4deeaaa6-1e01-4066-ad24-1c8f7e30df0a · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Safety in Large Reasoning Models: A Survey Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.070174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.070174Z digest=sha256:c930991304b6fc89c1e143849241ec102a13f1f4e4cd0f922ce66fa9e5df3c42

Observation 0605da23-8e2d-469d-88f9-c43235685d5f · outbound

This paper cites Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations.

Safety in Large Reasoning Models: A Survey Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:48.079314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:48.079314Z digest=sha256:fa10225b812eda4f55bba38e0b9abf0bf627f32263881a16e728fd36a96f0941

Pith citing papers

Observation ec2decda-44df-4610-84e2-d35b191b712e · inbound

Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction cites this paper.

Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction Safety in Large Reasoning Models: A Survey

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:11:57.987032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T19:10:55.009810Z digest=sha256:17b705e9f46fa3b84cfe8c6be46906a4d22e8f4ca8f387819f2c53f156318e92

Observation 9f01816e-fd69-4c38-a806-a699df400b22 · inbound

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework cites this paper.

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework Safety in Large Reasoning Models: A Survey

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.981064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.981064Z digest=sha256:bb517b5ba2f0550496463f4098adfb108d4e421188508231353aad5b2782cd18

Observation cd40c532-af5d-4742-ad30-1ddb0e71200c · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety in Large Reasoning Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.771889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.771889Z digest=sha256:5a01e8e868c836324deac3cb0a4065799419c7f76b3bcea87d5900b2fb706384

Observation a80fabca-9bde-4030-9093-a4157f3c9600 · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Safety in Large Reasoning Models: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.771631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.771631Z digest=sha256:64d50dd3d0fccbfa4ec04cde703d8a091a44e86f7bb6a64af6b92eea1190d855

Observation 4ec67418-7368-4711-ae4c-87b048c4808a · inbound

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents cites this paper.

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents Safety in Large Reasoning Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:18.251385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:18.251385Z digest=sha256:8fcf32b3c8dca4b30e1306def1f147595de54bc7b1bd3758d82d9afa0f753f26

Observation c2fb6f89-a64b-4de4-8f05-950781929a9e · inbound

Does More Inference-Time Compute Really Help Robustness? cites this paper.

Does More Inference-Time Compute Really Help Robustness? Safety in Large Reasoning Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.358800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.358800Z digest=sha256:f72d5c0d32b4c564b2a29f8b30cee2ddc54ccdc64d246f23d77e843127035167

Observation cb51eac4-6cb8-4d04-80be-8e44a7cef9d8 · inbound

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans? cites this paper.

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans? Safety in Large Reasoning Models: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:41.938259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:41.938259Z digest=sha256:f3fd91c43059717aed952c78927a6b1a8592971a87abb9eee5a0fe97902c9db0

Observation 3ad79d1a-3851-49ac-86ce-bf4c87be6a6e · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Safety in Large Reasoning Models: A Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.938124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:e906096911cb9ea6dd6817bb732bc0c3179c2887dba1874ab73e30069c0b1a96

Observation e7308aa5-db82-4b9c-8e55-a541a0873104 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Safety in Large Reasoning Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.351363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.351363Z digest=sha256:415c8b86d7039139d0fdd5ca19bdfd4aeb0026c943780baf2507e85e23451c6c

Observation dcc1b116-0f1d-4cc6-9788-e4250e40eb2a · inbound

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models cites this paper.

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models Safety in Large Reasoning Models: A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:05:55.672157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T05:02:52.176629Z digest=sha256:4d13015bddc2721ee9c4f9cee8dfd337f56e659c1f6553cde038cd05d246dd32

Observation 522417f4-1a78-4044-a568-0fd3d45fca66 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Safety in Large Reasoning Models: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.132863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.132863Z digest=sha256:c130afeb16cfcfb93db781986e7ccf4746eaf521f86a1c21e4028477a0398ac6

Observation 060770a7-9101-4996-9358-b12c61a49532 · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Safety in Large Reasoning Models: A Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:1865097aa7e0c07f40edeb80473ed63fc0e7cd791d2eb72f939337a0bb5b4ddb

Observation 975272f2-bc84-4b63-bf2a-e258ecb6aca2 · inbound

Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models cites this paper.

Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models Safety in Large Reasoning Models: A Survey

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:48:24.906392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T00:45:43.705160Z digest=sha256:d1c8bbec01e96efa42614542e224bf761625e691bb52e386e059dded3736f85f

Observation fccb413b-75e1-4a2b-a9c9-a8e50611f09a · inbound

DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection cites this paper.

DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection Safety in Large Reasoning Models: A Survey

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.379274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:54:24.013408Z digest=sha256:9246c0cdabab7da0c9ca14aeb8f3c1ed8399ecde6bac8ccf1e23469326201820

Observation 383970bc-d957-465a-a7c2-d419f55918ee · inbound

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing cites this paper.

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing Safety in Large Reasoning Models: A Survey

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.225748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T08:24:43.494005Z digest=sha256:1b3c9728f5b751a17cc8e7ff21d262eae5c0bd236651f06051c52c5d21c0a704

Observation 24a3568e-18dc-4507-aa29-21410f722811 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception Safety in Large Reasoning Models: A Survey

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.051370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:b00ff80e976b46803fa1f2b452c610cca7f24ea4790921212628c17bf4691d72

Observation cf3b9459-2add-4fe8-90e3-283ec66913e0 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Safety in Large Reasoning Models: A Survey

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:23.720034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:93c3b3fa4314b2ebcdbaaf914d2234bf6c8dc56b486dc935888ab1018afd4060

Observation e61e34aa-834a-4a49-ba36-142d27f82b0b · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Safety in Large Reasoning Models: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.580263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:b2e1283b4f54a012d964e8e620a85145e0c4d07e7e8fc3d97bc078deb8f79757

Observation 61f9dc42-46b5-4970-9d3c-59bc007c3153 · inbound

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows cites this paper.

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows Safety in Large Reasoning Models: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:06:20.139263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:06:20.139263Z digest=sha256:69b605349357c58c7db90e7a637b94b9cf803a55dc69ba0e41c2801f97fa1f12